← Insights Weekoverzicht 3 August 2026 14 min Written with AI assistance

The Plumbing Tell

This week the tells were in the plumbing — safety layers, scanners, batteries, and org charts made of agents.

Ruben Horbach Ruben Horbach Co-founder

In short

  • NVIDIA shipped Halos, a robot safety layer, before most robots have jobs — certification precedes scale.
  • Midjourney self-funded an ultrasonic CT scanner, showing a new model for building frontier hardware.
  • Models are becoming API front doors for multi-agent systems (planner/builder/judge org charts).
  • Chinese AI is compounding: token throughput doubled in a month, with revenue and rankings following.
  • Kyoto's cherry blossoms bloom weeks earlier — signals appear in the plumbing before the story.

Most weeks the AI news is a pile. A model here, a robot demo there, a jobs stat, a grid record. Each gets a moment and scrolls away before you can see how it fits.

This week the pattern was quieter than the headlines, and it is one I have started calling the plumbing tell: infrastructure keeps showing up before the thing it is built to support. Safety systems before the robots have jobs. Scanners funded by an image company. Org charts made of agents before anyone reorganises. If you watch the plumbing you can usually see the deployment curve coming.

So here is the week read from underneath, and then seven things I could not fit.

Safety infrastructure is a tell

NVIDIA shipped a full-stack safety system for robots before most robots have a job. It is called Halos for Robotics, and according to NVIDIA it rests on more than 18,600 engineering-years of autonomous-vehicle safety work, repackaged into a certified safety layer for machines that move in the same space as people. That engineering-year figure is the company's own accounting of its own history, so I would treat it as a claim about scale rather than a measurement.

Here is why that is worth a second look. Safety paperwork shows up just ahead of scale and never after it. Cars got crash standards before the highways filled up, and aviation got its certification stacks before the sky got crowded. The certification tends to arrive a step before the deployment curve does.

So when the company that sells the compute for humanoids also starts selling the safety layer for them, I would read the sequence rather than the product. The machine that keeps robots from hurting people is being built while the robots are still learning to fold laundry. That order of operations is the signal. The announcement is not.

The obvious objection is that a safety layer is also a moat, and I think that objection is correct rather than competing. Certification is expensive to build and expensive to displace, so whoever defines the reference stack early gets to sit underneath everyone else's product for a decade. Both readings point the same way: you do not spend 18,600 engineering-years on a layer for a market you expect to stay small.

What I would actually watch is who adopts it. A safety stack only becomes infrastructure when the second and third robot company build on it instead of writing their own. As far as I can tell that has not happened yet. Until it does, I would call this a bet with a very confident balance sheet behind it rather than a settled standard.

Infrastructure keeps showing up before the thing it is built to support.

The scanner that came from an image company

Zero VC funding. Around $500M in revenue. Fewer than 200 people. And they just reinvented the full-body scan.

The company is Midjourney, the one you know for generating images. It took no outside money, funded itself on that image product, and quietly spent the profits building hardware. The first result is a scanner you lower into water for about 60 seconds while it maps the tissue inside you down to sub-millimetre detail, with no radiation, no magnets and no X-rays. Founder David Holz calls it ultrasonic CT: as detailed as an MRI, as easy as a trip to the spa. On some tissue measures it already matches MRI on day one, at a fraction of the cost. Those are the company's own claims on a first-generation device, so I would wait for independent comparisons before treating the MRI parity as settled.

The machine is impressive and the structure behind it is the actual story. A profitable software product became the R&D budget for a medical device, with no term sheet, no board and no permission. That is a new shape for how frontier hardware gets built, and it runs on the same principle as the safety layer: build the capability before the market asks for it.

It also has a wall in front of it that no amount of revenue removes. A diagnostic device meets regulators, and regulators are indifferent to how elegantly it was funded. The path from a working scanner to a machine a hospital may point at a patient runs through trials measured in years. What the funding structure buys is probably not speed through that process. My guess is that it buys something else: the ability to sit inside that process without a board asking quarterly when the money comes back. Plenty of promising medical devices seem to die of impatience rather than of physics, so that may well be the constraint that actually matters.

The model is becoming a wrapper

Sakana AI reports that its Fugu-Ultra lands in the same performance range as Claude Fable and Mythos. That was the headline, and the mechanism underneath it is the part worth your attention. Worth noting that these are the lab's own benchmark placements, which is normal for a launch and still not the same as an independent evaluation.

Fugu is a multi-agent orchestration system sold as an ordinary model API endpoint. You call one URL and a team of agents does the work behind it, and to your code it looks like any other model. Anthropic's own engineers gave the same idea a face this week. They build apps with three agents in a loop: one planning, one building, one judging. On their account a full app went from nothing to shipped in about 40 minutes, with a human sitting on top of the pipeline rather than typing inside it.

The speed is not the point. What you are looking at is an org chart. Planner, builder, judge: three roles that used to live inside one developer's head, now split across agents that hand work to each other while the human runs the review layer and decides what good looks like.

The cost of that arrangement is real, and the endpoint hides most of it. An orchestrated system fails in ways a single model does not. Agents disagree. A planner commits to a bad decomposition and the builder faithfully executes it. The judge approves work that is locally correct and globally wrong. You also pay for every internal turn, so I would expect the same answer to cost several times what a direct call would. None of that shows up in the benchmark score, because the benchmark only sees the output.

Once the endpoint is the only thing you touch, "which model" stops being the right question and what the system does when you are not looking becomes the whole of it. The model is quietly turning into the front door of a coordinated system, and the buyer never has to know the difference. That is convenient right up to the first time you need to explain to a regulator what actually made a decision.

China stopped catching up and started compounding

Chinese AI models processed 98 trillion tokens in June, against 46 trillion in May. That is a 113% rise in 30 days, roughly a doubling inside a single month. Around the same time DeepSeek is reported to be closing in on $500 million in annualised revenue and lining up an IPO, and Kimi-K3 took the top spot on the frontend code arena with a 48-point lead over Claude Fable 5.

For two years the story about Chinese AI was catching up on benchmarks, and that story is running out of road. Most of the debate still fixes on leaderboard rank, when tokens processed reflects real work: people and systems actually running these models at scale. The FT keeps documenting how many apps you already use run on Chinese models under the hood.

Now the number that cuts the other way, and it is a big one. On the cluster-level estimates I have seen, the US holds around 75% of the world's GPU-cluster compute and China around 15%: roughly a 5x gap in raw training capacity. Those shares depend heavily on what counts as a cluster, so read them as an order of magnitude rather than a precise split. Serving models cheaply at volume and training the next frontier model are different races. China appears to be winning the first while losing the second by a wide margin. Efficiency compounds when the frontier stands still; when it moves, capacity is what moves with it.

So put both together rather than picking one. Revenue funds the next model, ranking pulls in the developers, distribution locks in the users, and none of that closes a 5x compute gap on its own. My read is that the next stretch gets decided on infrastructure economics, meaning cost per token and who can serve it cheapest at volume, which is a race China is genuinely positioned to win. Whether that translates upward into frontier training is the open question, and a single month of doubling is a short base to extrapolate from. I would want another quarter before calling any of it a trend.

The curve you can feel

Kyoto's cherry trees bloom two to three weeks earlier than they did for centuries. The 30-year average peak blossom date has slid from mid-April, around the 15th to 17th, to late March by the mid-2020s. No model, no projection, no argument about methodology. Just a date, hand-recorded across generations of festivals, bending steadily in one direction.

This belongs in a week about infrastructure because it is the same lesson from a different domain. The trees are reading conditions and writing the answer down, quietly, ahead of the debate. Climate usually reaches us as a chart of anomalies most people cannot feel, and this arrives as a calendar: the day a city has gathered under the blossoms for as long as anyone remembers, arriving earlier each generation.

The honest caveat is that Kyoto is a city, and cities warm for reasons that have nothing to do with the atmosphere. Concrete holds heat, and a record kept in one place across centuries of urban growth is measuring both effects at once. Researchers who work with this series say they adjust for the urban component and that the trend survives the adjustment. I am taking that on trust rather than checking it myself, and it is the part that makes the record useful rather than merely charming.

What I take from it is a habit rather than a fact. The most reliable indicators tend to be the ones nobody designed as indicators: a blossom date recorded for a festival, a token count logged for billing, a certification filed for liability. Instruments built to persuade you are built by someone with an interest in the reading. Instruments built for some other purpose entirely are the ones I trust.

The through-line is that every big shift this week announced itself in the boring layer first: the safety cert, the funding structure, the API endpoint, the token count, the calendar. The headline follows later. And the reason the plumbing runs ahead is not that anyone is hiding the story, but that certification, funding structures and billing meters all have to exist before the thing they support can be sold at scale, which means the people building them are committing to a forecast months before the rest of us get to read about it. That commitment is the tell. In my experience reading it is worth more than being early.

Seven things I could not fit

From half a million to nine million in five months. OpenAI's Codex went from under 500,000 weekly users in early February to 9 million by mid-July. Layoff headlines get the attention, but a curve like this is the thing that actually reshapes a job: developers quietly making an agent part of the default workflow until it is simply how the work happens. Most of them still run one agent at a time, which is worth remembering before anyone extrapolates.

The fastest-growing group of AI users is over 65. The most pessimistic is under 30. Same dataset, two lines pulling apart: every age bracket used more AI over two years, the biggest percentage jump came from the 65-plus group, and 48% of 18-to-29-year-olds expect negative consequences against 39%, 38% and 35% for the older brackets. The generation that grew up digital is the one bracing for impact, and I would guess that is about entry-level work rather than about understanding the technology.

Holding several objects in mid-air, with nothing touching them. ETH Zurich's Multi-Scale Robotics Lab is steering multiple items through the air at once, contactlessly. Most of robotics right now is a fight over hands — tendons, tactile skin, a grip firm enough for a wrench and gentle enough for an egg. Contactless manipulation walks around that problem rather than through it, which is either a narrow laboratory trick or a different branch of the tree entirely, and it is too early to say which.

Hearing a fan fail before the sensor does. Spending time with data-centre operators this past month, the bottleneck in physical AI was not where I expected. Robots copy motion well: give them demonstrations and they will match the walk, the reach, the torque. What they cannot copy is the operator who hears that a pump's pitch is slightly off, or knows which rack runs hot in summer. That judgment was never written down anywhere, which is precisely why it is hard to transfer.

A thousand strike drones a month, from one European factory. Helsing finished a plant in southern Germany with capacity for 1,000 HX-2 drones monthly, on top of 4,000 already delivered to Ukraine and 6,000 more committed. The Ukrainian drone story has mostly been improvisation — workshops, garages, FPV rigs taped together. Serial production on European soil owned by a European company is a different thing, and it says more about rearmament than another debate about percentages of GDP.

Storage caught up with the sun. The share of new daily solar generation that new Chinese storage can time-shift climbed from 6% in 2022 to 21% in 2025, more than a tripling in three years. Everyone tracks how much solar gets installed and almost nobody tracks whether the power is usable after dark. The standard argument against renewables assumed that gap stayed open.

The builder and the intruder run on the same engine. OpenAI's own system accidentally breached Hugging Face; in the same stretch the security firm Irregular published an assessment of GPT-5.6 Sol against offensive-security benchmarks, and Okta shipped a risk-assessment tool for enterprise agents. Three signals from one corner of the industry pointing the same way: the capability that writes and ships your code is the capability that can probe a system and move data out of it.

This is the connected version of the week. If it is useful, it lands in your feed every Monday.

Ruben Horbach

Ruben Horbach

Co-founder · Back From the Future

Ruben researches how organisations adopt AI meaningfully — not as technology, but as a change in work and people. He builds the agent infrastructure behind BFF and speaks about the near future of work.

Translate this to your situation?

Book a conversation — we're happy to think along about what this means for you.