Teng Li · ML Engineer

Instinct moves $1B a year through SMS and email. Its bottleneck is the one I've been measuring.

Instinct raised a $1B Series C this week at a $10B valuation. Fourteen employees. Invite-only. Over $1B in annualized transaction volume, half of it travel. Zero marketing spend.

The valuation isn't the interesting part. The interface is: there is no app. You text it, you call it, you email it. It has its own phone number and its own computer, and it calls you when something is urgent.

That one decision pushes every hard problem into a single layer — the place where the agent actually touches the world. Which is the layer I've spent the last two months measuring, from the opposite end.

The number nobody is quoting

Buried in the founder's interview is a metric that tells you more about this industry than any benchmark I've seen this year:

After three weeks on the product, 40% of users have shared a personal credit card. Once a user shares at least one piece of sensitive information, retention is 80%.

He treats time to first credit card as his proxy for trust.

Read that again. The growth curve of a $10B company is a permissions curve. Nobody is churning because the model reasoned poorly. They're churning because they never climbed high enough up the trust ladder for the thing to be useful.

This is the same finding from the other side

Two months ago I lint-scanned 36 popular MCP servers and found that roughly a third are effectively unusable by an agent. Spec-compliant. Well-built. Still broken in practice — ambiguous names, undescribed parameters, errors that tell the model nothing.

Instinct is measuring how fast a user will grant an agent permission to act. I was measuring whether, once granted, the agent can act competently.

Different ends of the same pipe. Neither end is the model.

1. A thinner interface means more integration work, not less

"No app" sounds like less engineering. It is the opposite.

When your entire surface is SMS and a phone call, there is no UI to disambiguate intent, no dropdown to constrain a parameter, no confirmation modal to hide behind. All of that complexity doesn't disappear — it relocates into tool design, into the permission model, and into how the agent recovers when a call fails.

The teams shipping the thinnest interfaces are carrying the heaviest integration burden. Worth remembering the next time a clean demo makes the underlying problem look solved.

2. Trust is a ladder, and every rung is a product decision

Read-only. Then write. Then spend my money.

Most agent products treat permissions as a security checkbox bolted on before launch. Instinct's numbers suggest it's the core growth mechanic: each rung a user climbs unlocks a materially more valuable product, and the climb takes weeks.

If that's right, the interesting questions stop being "is this secure" and become:

OpenAI's recently open-sourced Codex harness gets this right, whatever you think of the product. Approval requests and interruptibility aren't bolted on — they're first-class citizens of the execution loop, sitting alongside tool invocation and event streaming.

Approval is the loop. Most agent frameworks still treat it as middleware.

3. Proactive agents have a different compute shape

Instinct's founder says he spends 40% of his time on compute. Not because of user count — because of shape. The agent wakes at 6am to prepare your day. That background load simply does not exist in a request-response product. He reports 3–8x efficiency from splitting background work from real-time interaction in a custom inference deployment.

In the same month, OpenAI reported that harness-level changes alone — retained reasoning and context compaction — moved GPT-5.6 Sol from 13.3% to 38.3% on ARC-AGI-3 while cutting output tokens 6x.

Same model. Different scaffolding. Two very different bills.

We spent two years optimizing prompts. The next round of gains looks like it lives in the execution loop: what you load into context, when you drop it, and which work runs on the user's schedule instead of the request's.

The common thread

There's a background assumption that this race gets won by whoever has the best model. The evidence keeps pointing somewhere less glamorous.

A fourteen-person company is moving a billion dollars a year through SMS and email, and its founder's two stated bottlenecks are how fast users will trust it and how much compute the background work eats. Neither is a modeling problem. Both live in the layer between the model and the world — connectors, auth, tool schemas, permission ladders, approval UX, retry semantics, context budgets.

That layer doesn't demo well. It shows up in your error rate and your invoice.

It's also, as far as I can tell, where the next few years of differentiation actually are.


Next up: mcpgrade currently scores whether a tool is usable by an agent. It says nothing about what that tool costs — how many tokens a typical call drags into context, and how much of that survives to the next turn. Given the numbers above, that might be the more interesting question. I'm adding a context-cost dimension and re-running the 36.


← All writing