Blog
How we build the products, what the numbers actually look like, and what we got wrong.
Anna: the agent we built by taking things away
Anna answers a trade business's WhatsApp day and night. The interesting part is not what she does. It is the list of things she is structurally incapable of doing, and why that list is the product.
The $1.20 Problem: Why Agentic AI Workflows Cost 5x to 30x More Than Chat
Gartner measured a 5x to 30x token multiplier for agentic workflows over chat. Here is where the multiplier comes from and how to manage it before your bill does.
Otto: what a local-first meeting agent looks like at month three
Otto is a macOS app that captures your calls, transcribes them on your own machine, and briefs you before your next meeting. It is in private beta with one pilot user, and this post is a straight account of what works and what does not.
A Practical Guide to LLM Model Routing and Cost
Model routing sends simple queries to cheap models and complex ones to expensive ones. Here is how to decide which LLM tier to use based on cost and capability, with worked examples from the current pricing landscape.
Cost: prove a cheaper model is safe, then keep the saving
Everyone knows they are overpaying for inference. Almost nobody downgrades a model, because nobody can prove the cheaper one is good enough. Cost proves it first, on your own traffic.
How to Reduce Your OpenAI API Costs
Six OpenAI-specific techniques for cutting your GPT bill: prompt caching, model tiering, structured output discipline, tool call costing, prompt compression, and batch processing.
The Compute Trap: Why AI Startups Burn Through Seed Capital on Inference
One analysis puts AI startups at 40 to 60% of cash spent on infrastructure before the product is proven. Here is how to track LLM cost per feature, per user, and per business action before cash runs out.
Sage now joins your Microsoft Teams calls
No new setup, no separate login, no Teams-specific mode. If a Teams meeting is on your Google Calendar, Sage shows up - same document, same email, same way.
Why AI Meeting Agents Are Replacing Meeting Notes
Passive transcription tools capture words. AI meeting agents capture intent, act on it in real time, and deliver outcomes before the call ends.
Meet the Line-Up: One Idea, Measured Three Ways
Cost and Watch measure two kinds of agent drift and Eval, in development, will measure the third. Anna and Otto sell to a different buyer. Sage, Cole and Lex are still reachable but no longer where the work goes. Here is how the line-up fits together, and why it shrank.
The Hidden Cost of Meetings: How Unstructured Follow-Up Kills Momentum
Decisions made in a meeting decay fast: if nobody writes down who owns what before the next call, the commitment quietly stops existing. Here is what a structured meeting document changes, and why the delivery speed matters more than the format.