Four separate labs cut the cost of running an agent to near zero inside five weeks, and one of them cut it to actually zero on hardware you already own. None of them made the web easier to read. When compute stops being the constraint, whatever is left becomes the constraint.
For two years, the honest reason most agent ideas stayed on the whiteboard was arithmetic. The workflow made sense. The token bill did not. You could describe an agent that watched a market all day, and then you worked out what a million tokens an hour costs and quietly built a dashboard instead.
That constraint broke in about five weeks. Here is what happened, and then the part almost nobody is saying out loud.
Four events, one direction
30 July. OpenAI cut GPT-5.6 Luna by 80%. Input went from $1 to $0.20 per million tokens, output from $6 to $1.20. Terra took a 20% cut. Sol was left alone. The cut landed three weeks after the GPT-5.6 family launched, which is not the behaviour of a company setting prices from a cost model. It is the behaviour of a company reading a competitor's price list. CNBC's own reporting points at why: Chinese models had taken 46% of US enterprise token usage on OpenRouter, at times running ahead of US models outright.
31 July. DeepSeek put the V4-Flash API into public beta, with the agent benchmarks rebuilt rather than the model. Same architecture, same size, redone post-training, and scores that now pass its own larger V4-Pro preview on agent work. It also picked up native Responses API support. The cheap model got better at being an agent, specifically, and it did it without getting bigger.
3 August. Alibaba announced Qwen3.8-Max, a 2.4 trillion parameter model, at launch pricing of $2.00 per million input tokens and $6.00 output, with implicit caching at $0.25. Then it said the weights are going open, which would be the first time a Qwen-Max-class model has been open-sourced, alongside Qwen3.8-27B. A frontier-class model priced like a mid-tier one, with a free version arriving behind it.
4 August. Liquid AI released LFM2.5-2.6B, and this is the one worth sitting with. It is a 2.6 billion parameter agentic model that plans, calls tools, and works through multi-step tasks entirely on a device you already own. 128K context. Under 2.5 GB of memory. Around 30 tokens per second on a phone, 220 on an M5 Max. On Liquid AI's published benchmarks it scores 77.83 on ToolSandbox and 85.49 on IFStruct, ahead of Qwen3.5-9B on both, at a fraction of the size.
The data never leaves the device. The marginal cost of a run is nothing.
What gets cheap gets used
This is the oldest pattern in industrial history and it never stops being true. When something expensive becomes cheap, consumption does not stay flat and save you money. It expands until the cheap thing is everywhere.
Agents are already following the curve. When a query costs real money, you write an agent that runs when you ask it to. When a query costs a fraction of a cent, you write one that runs on a schedule. When it costs nothing at all, because it is running on a laptop that is switched on anyway, you write one that never stops. Liquid AI's own framing is that local agents can be parallelised across hardware you own and burn through millions of tokens on background tasks at no marginal cost.
So the number of agent runs is about to go up by an amount that is difficult to picture. Not because anyone decided the agent economy should scale, but because the thing that was capping it stopped capping it.
The bottleneck moved. It did not disappear
Here is the part that gets skipped, and it is worth being precise rather than tidy about it.
Cheap inference does not create demand for data. That would be too convenient an argument, and it would be wrong. What cheap inference does is remove compute from the list of things limiting an agent. And when you take the binding constraint out of a system, the system does not become unconstrained. It just finds the next constraint.
For anything that touches the outside world, the next constraint is obvious once you look for it. An agent with free thinking and no information is a very fast way of being confidently wrong. Models do not contain today's prices, today's listings, today's search results, or today's version of a page. They contain a compressed memory of a web that has since moved on.
That gap is not narrowing. It is widening, because the same period made every model cheaper and none of them fresher.
A phone cannot crawl the web
The on-device case makes this sharpest.
Put LFM2.5 on a laptop and you have a capable agent with zero inference cost and no route to the live web at any useful scale. It has one residential connection, no rotation, no geographic reach, and no way to check whether the page it just fetched is the page a real person in another country would have been shown. It can think all day for free. It cannot see.
And the web it is trying to see is closing. Automated traffic passed human traffic in June, which gave every site operator a reason to harden. Datacentre IP ranges get blocked outright. Requests get rate-limited. Sometimes a scraper is quietly served a stripped-down or cloaked page and never finds out it took away bad data, which is the failure mode that actually hurts, because it does not look like a failure.
So the shape of the next two years is already visible. Abundant, nearly free reasoning. A tightening supply of anything true and current to reason about.
What that means for a Beacon
Beacon is a node application you install on a device you already own, sharing a bounded slice of unused internet bandwidth. It runs natively on Windows, macOS, Linux, Android and iOS. Operators earn Teneo Protocol Points and Fragments for what they contribute, and claiming accumulated Fragments every eight hours activates a reward boost worth up to 3x. Every unique IP earns independently, so nodes stack across devices. One node on its own is unremarkable. The network of them is the point.
What that network provides is an honest view of the open web. Real residential and mobile connections, on real devices, in real places, seeing pages the way a local person sees them rather than the way a datacentre is permitted to. That is the difference between checking a price and checking the price a customer in that country is actually shown.
Everything that got cheaper in the last five weeks got cheaper because it could be copied. A model is weights. Weights can be distilled, quantised, open-sourced and undercut, and in the last fortnight all four of those happened to somebody. A price cut takes an afternoon.
Access to the live web cannot be copied. It is physical supply: devices, connections, geographies, uptime, and people choosing to opt in. Nobody ships that as an SDK, and no lab can price-cut its way into having it. That is why the buyers were already here before any of this happened. AI teams gathering data, market intelligence and e-commerce platforms tracking live pricing, ad verification companies confirming a campaign ran where it was paid to run. They are not buying a finished dataset. They are buying reliable access to the live web, and it is getting scarcer every quarter.
The other end of the same network
The agent layer sits on top of it. Hundreds of live agents on the Teneo Protocol, one MCP-compatible CLI, pay per call in USDC over x402. No API key, no subscription, no seat. Market data, on-chain analytics, public X data, all queryable from whatever stack you already run.
That is the pairing that matters as inference goes to zero. An agent that costs nothing to run still has to pay for something worth knowing, and it should be able to buy that in the same call it pays for everything else, from a network that can actually reach the place the answer lives.
The models got cheap. Being right did not.
Two ways in. Install Beacon and contribute bandwidth from a device you already own. Or install the CLI and start calling agents. Both start at teneo-protocol.ai.
Key takeaways
- -Inference costs
- -On-device AI
- -Agent economy
- -Beacon
- -Teneo CLI



