Claude Sonnet 4.6 achieves new benchmarks /// OpenClaw hits 50K stars /// Vibe Coding surges 340% /// Claude Sonnet 4.6 achieves new benchmarks /// OpenClaw hits 50K stars /// Vibe Coding surges 340% ///
Uncategorized

A scheduled task is not a returning user

Google is charging $99.99 a month for an agent that runs whether you watch it or not. Every retention metric built for humans counts that cron job as an engaged customer, and the bill arrives before the graph moves.

I have a job that runs at 6:40 every morning. It pulls the day’s research and drafts a post. Then it grades the draft against a checklist and drops the verdict into a Slack channel where I’m the only member.

One week in July I stopped reading it. Not deliberately. The thread was there each morning, I glanced at the first line, and then I went to do something with a deadline attached. The job never noticed. It ran on time and logged its cost, seven mornings out of seven. From the outside, that’s a week of healthy daily usage by a user who had, franchement, checked out.

August put a price tag on that gap. Google is now selling Gemini Spark at $99.99 a month, an agent that will phone a store and place an order on your behalf, and Anthropic is selling Claude Cowork at $20. Assistant is gone from Android too, replaced system-wide by Gemini. The agent isn’t an app you choose to open now, it’s the layer underneath. One figure making the rounds puts active consumer agent use at 25%, and I can’t trace it past a news roundup, so treat that as vibes with a decimal point on it. The subscriptions are real, though. And they will be measured with tools built for a world where a login meant a person.

Where the churn hides

Retention math counts events. Sessions, active days, tasks completed, weekly actives. Every one of those metrics assumes a human on the other end deciding to come back, because for twenty years that assumption was free.

On the vendor side, Userpilot’s 2026 write-ups on adoption metrics for humans versus AI agents describe the break cleanly: agent activity continues on its schedule regardless of human engagement, so an account reads as retained after the person has quietly stopped logging in. They pair it with what they describe as the agent churn fingerprint, heavy abandonment inside the first 30 days plus shallow feature use. That’s a product-analytics company writing about the value of product analytics, so label it accordingly. The mechanism doesn’t need their survey to be true. It needs a cron job and a dashboard, and both exist.

Compare it to the failure mode we already know. A BI dashboard nobody opens goes quiet. Sessions fall, the graph droops, and somebody on the customer success team sees it in week three and books a call. The decay is the warning. An unsupervised agent produces the opposite shape. It keeps writing rows, the usage line stays flat and healthy, and the first real signal is the cancellation itself. You lose the early warning precisely because the product got better at working without you.

Then there’s the cost side, which compounds in the same direction. OpenAI cut GPT-5.6 Luna to $0.20 per million input tokens, down 80%. Do the division: $99.99 buys roughly 500 million input tokens at that rate. Spark doesn’t run on Luna, output tokens cost more than input, and neither company publishes its per-task consumption, so that number is a floor for what the cheap end of the market can burn and not a margin estimate. What it tells you is the shape. An always-on agent’s spend scales with its schedule. It does not scale with whether anyone read the output.

The 87% problem

The enterprise numbers have the same crack running through them. NVIDIA’s 2026 state of AI report has 64% of organizations deploying AI in operations rather than piloting, 87% reporting cost reductions, and 88% reporting revenue increases. NVIDIA sells the hardware underneath all of it, which doesn’t make the survey wrong, it makes it labelled. More to the point, those are self-reports. Call it a poll of people who already bought the thing, asked whether the thing worked.

Account-level metrics make the masking worse at that scale, not better. A company with 400 seats and a nightly agent workflow rolls all of that into one healthy-looking account. Scheduled jobs keep firing across seats where the humans have drifted off. Nobody’s lying. The number just answers a question about machine activity while everyone reads it as a question about people.

Bon. Another force is arriving at the same time, from a direction nobody planned for. Since August 2, 2026, the EU AI Act’s transparency obligations have been legally binding. Tell users when they’re dealing with an AI system. Watermark generated content. Fines run up to €15 million or 3% of global turnover. Everyone’s reading that as a compliance chore about labels. The interesting part is that mandatory disclosure and mandatory watermarking together produce a record of which outputs were machine-made, which is most of the instrumentation you’d need to separate human work from agent work. And disclosure cuts the other way too, since 81% of consumers say they fear AI access to their data. Every honest label is also a small reminder to cancel. It’s not because a notice is required that it’s neutral.

Here’s the strongest case against all of this. I’m predicting, not measuring. There’s no public cohort data on consumer agent retention past 90 days, from Google or Anthropic or anyone else, which means my claim about the flat line before the cliff is a mechanism story with no curve attached to it. Google’s analytics people are not amateurs, and I’d bet their internal dashboards already separate human-initiated from scheduled sessions. I can’t see those dashboards. My whole argument could be a description of a problem that competent teams solved eighteen months ago and never wrote about. Me, I have a bias toward stories where the metric is the villain, and that bias has cost me before.

What would move me: a retention curve from either company that splits human-initiated sessions from agent-initiated ones at 90 and 180 days. Not aggregate actives. If those two lines track each other, I’m wrong and the delegation model is healthier than I think. If they diverge, everything reported as adoption this year needs an asterisk.

Until then, if you run a product with agents in it, go split the metric yourself before the renewal cycle splits it for you. Count human touches, not events. Log the moment a person last read something the agent produced, and treat that timestamp as the retention number.

Mine would be one line of code. I haven’t written it, which tells you roughly how much I want to know the answer.

Dominic Plouffe

Staff writer at Neural Pulse.