← All posts
# AI 6 min read

Owner Agent, Part 1

Building Owner Agent, starting with support.

Daniel Ternyak
Daniel Ternyak
Head of AI Agents
224k+historical support cases mined
~2,000discrete operations they reduce to
~1,000evals stubbed by hand so far
80%of cases arrive by phone

Owner Agent needs breadth to handle both the common operational tasks and the more advanced questions.

Common operational tasks: Some customers won’t use the Owner dashboard, or any dashboard. They will call support even to change their restaurant’s hours, which takes 3 clicks in our dashboard. They’ll still call. We want to be super available to them with no hold time.

Common tasks: one sentence, one approval card

More advanced questions: Some customers run dozens of locations. They want to understand trends, like what's moving on their menus. They want to know how many corporate customers they serve, and keep track of what their competitors are doing. There’s room for sophistication in helping them grow their business.

Owner Agent answering a question about a restaurant's business

In the medium term, Owner Agent will also become proactive, volunteering to take actions to grow each customer’s business (e.g. create a coupon code in time for National Taco Day to drive app orders). Owner Agent will be the gateway to the entire breadth of the platform.

So we had to decide where to start - what to prioritize, which classes of problems to solve now versus later.

Why we decided to start with support

We've started by framing Owner Agent’s first big "customer" as our own customer-facing teams: 100+ users, thousands of conversations per week.

They use Owner Agent to help with time-consuming support tasks - for example, setting up a new menu. In the past, our reps had to manually review a Google Drive with hundreds of food images, half of them unlabeled, trying to eyeball which image corresponded to which menu item. That pain is gone. Owner Agent has vision as one of its modalities; it offers its best guess to match menu items to photos. It also does the previously time-intensive work of deconstructing a menu PDF into items, modifiers, and assets.

Owner Agent breaks a menu PDF into items and modifiers, then matches unlabeled photos to items

Soon Owner Agent will be customer-facing. This isn't because our reps aren't great: in fact customer satisfaction exceeds 90%, and customers often mention how much they love our reps. But at peak times, callers can still wait on hold for 15 to 30 minutes. You shouldn't wait 30 minutes to change your hours just because the dashboard feels intimidating. Ideally Owner Agent will cut into those wait times.

3 primitives

3 primitives run through the rest of this post. Our team looks at them as a hierarchy:

A request flows from its channel, through a tool, into a capability proven by evals
01Channel
How the customer talks to us: the support line, email, or in-product through Owner Agent in the dashboard. 80% of cases today come through the phone. This matters because our ability to represent an action, the tool call, depends on the channel. A complicated menu change isn't safe to express over voice. Voice is lossy.
02Tool
The ability to perform a narrow operation: change a menu price, update a photo. We wrap existing service methods in our codebase. The codebase is full of functions that change hours or refund a customer. A tool is the bridge from that function to something an agent can activate.
03Capability
The underlying set of tools and prompts that gets an agent to an outcome. Evals are how we show we have a capability - that we can handle a type of request reliably. We’ll come back to this.

For example, suppose the customer says:

Remove the pollo taco from the menu at Miramar, add tomatillo as a salsa option there, and raise the price of guacamole in Atlanta by $1.

Even our human support reps would write up an email to confirm. The channel shapes how a tool call gets represented, and each channel has its own human in the loop.

It’s a hierarchy because, for instance, the channel is upstream of how Owner Agent gets approval to call a tool.

Tool call approvals

The approval is the agent saying it's about to make an operation and needs to show you first, then deciding which channel that comes into: UI or SMS. For example, imagine you call in over the phone wanting to make a meaningful change to your menu, including changing modifiers. We'll repeat the whole set of operations back to you: the intake, then the confirmation, out loud - but there's still room for error. We convert the set of operations into plain language and send it over SMS, confirming all the changes. So when you call in, you become the human in the loop. :)

A spoken request, confirmed in plain language over SMS before anything changes

That isn't an industry convention. But it's what our customers need, because we can't count on them to have our mobile app.

But say you do have the app. In that case, we can link you to your account and trigger a push notification that shows you exactly where to find that feature.

A push notification linking the owner to the exact feature in the app
With the app installed, a push notification links to the exact screen

This goes for our reps too. They can say: “I'm happy to do this for you. But if you'd rather not talk to me, I just dropped a push notification into your phone.”

It’s educating the user inside a flow that was going to happen anyway, showing a better way to get there next time.

From cases to evals

We have converted our historical support cases into a provable mechanism to demonstrate that the agent has the capabilities to impact these cases. That mechanism is our eval dash.

With help from our BizOps team, we’ve mined Owner's 224,000+ historical support cases from Snowflake and clustered them into a small set of scorable capabilities.

224,000 cases collapse into about 2,000 operations, each backed by evals
Eval dashboard showing support case categories
The eval dash, top-level view

The top-level view covers about 25% of our case categories. Take menu updates: roughly 10,000 of our 224,000 cases, about 5%, came from someone trying to change a menu. Each case carries 3 levels of categorization: primary, secondary, tertiary. We counted 414 distinct reasons at last check.

Eval dashboard drilling into a case family ranked by time to close
Case families ranked by support time, not volume

We rank the work by support time, not case count. A family can match another in volume but take 10x the number of hours. So we show both and prioritize by time. That's why the target leads with time-to-close. We want the agent to absorb about 50% of the time our cases take to resolve.

Compressing 224,000 cases into 2,000 capabilities

The cases overlap more than they look. Say one case needs 3 operations and another needs 4. If 3 of them match, the pair shares only 4 distinct operations. The count collapses. Today it looks like 224,000 cases reduce to around 2,000 discrete operations. Build those 2,000, and we cover the next 1,000 cases that lean on them.

One distinction: a capability isn't the same as a solved case. Hand the agent a case with 70 moving parts and it may need to break it up, or punt, as a human might too. So we track 2 axes: the capability on one, reasoning through a specific case on the other.

I've stubbed about 1,000 evals by hand. Next I wire them into the codebase and run them in Buildkite. Then we run the agent against them. A second LLM grades each run against the expected output and tools, and we watch the trend. In production, a separate LLM checks every response against the tools the agent called. If the agent says "I checked your account" but never called the lookup, we catch it.

If the agent passes the evals tied to a set of cases, we can say we solve those cases.

Coming up in later editions

  • Escalation. What if there’s a 2-3 week support engagement with 16 turns? What if the agent can handle 7 turns, but escalates on the 8th? We will build for those possibilities.
  • Observability. Customer-facing teams need to be able to observe and understand what the customer wants the agent to do and what it did. This infrastructure will be unified so that SMS, email, and every other channel will be in one timeline.