Nicholas Hong← All work

Prototype · built hands-on

A multi-agent travel assistant, built to find out

A working prototype of the whole product loop, built alongside my PM role and kept separate from the production platform. It carried live investor and partner demos, put product ideas in front of real users early, and gave engineering evidence to make their calls with.

Role
Product manager and builder, alongside the PM role
When
2025 — 2026
Context
Sidekick Labs, prototype and demo platform
Obvious requests skip the coordinator. Ambiguous ones go through it, with the conversation's state. Every answer is one typed card, drawn richly on the phone and glanceably on the glasses.

Why build it

Production engineering had its own roadmap, and some questions couldn’t wait for it. We needed three things that slides can’t provide:

  • a product investors and partners could actually use
  • somewhere to test ideas with real people on-site
  • evidence for open architecture questions: how requests get routed, where state lives, and what latency is achievable

So I built the full loop myself, on a separate stack so it never competed with production for engineering time: an AI backend, an Android phone app, and a heads-up display app for glasses. Investors, partners and people on-site used it live, so a demo still had to behave like a real product.

Architecture

A coordinator and specialists

A stateful coordinator agent receives every request and delegates to specialist agents. They cover:

  • recommendations
  • restaurant and sightseeing bookings
  • rides and transit directions
  • calendar and reminders
  • visual question-answering about what the user is looking at

Later additions included messaging and an agent whose behaviour is defined as data rather than code.

Splitting work across agents isn’t free. Every hand-off is another model round-trip, which matters on glasses. The split was worth it because each specialist could be tested and changed on its own, but most of the later architecture work went into clawing that latency back.

State, because stateless agents lie

The first version was stateless, and it failed in instructive ways. Agents confirmed bookings they had never made, and skipped steps in multi-step workflows.

The fix was to persist conversation state for the coordinator and every specialist, and rebuild it from the database if it expires. That also fixed follow-ups like “book that one”, where the agent needs to remember what “that one” refers to.

Routing in two tiers

  • Tier one is deterministic. If the context already makes the answer obvious (for example, a photo arriving in a known setting), the request skips the coordinator and goes straight to the right specialist. One model round-trip instead of two.
  • Tier two is the model choosing tools. A validator checks the sequence of tool calls afterwards and retries if it’s wrong.

A tracker running alongside the conversation

Chat history alone can’t reliably answer “what’s happening with my booking?” So a lightweight model watches the conversation and keeps a structured record of each task in flight: booking, directions, reminder, and so on.

Each record has a lifecycle: active, suspended, completed or abandoned. It drives a persistent “what you have going on” tray in the apps.

One response, two renderings

Every response carries a typed card, such as a booking, directions, a recommendation or a ride. The phone and the glasses each render it in their own way: rich on the phone, glanceable on the heads-up display.

Decisions worth explaining

  • Latency is the product. On glasses, a slow answer is a wrong answer. The architecture was reshaped repeatedly to remove model round-trips:
    • deterministic routing
    • a fast pre-identification step that skips the large model when confidence is high
    • a direct mode for visual questions
  • Keep data away from the model where the data matters. Cards that show numbers bypass the model completely. If a model copies figures into a card, the card doesn’t just fail to appear; it can show the wrong numbers.
  • Fail toward honesty. When photo recognition is unreliable, the assistant asks the user to describe the object, or to photograph the information board, instead of guessing.
  • Test ideas cheaply. Capabilities defined as data, with a web interface to create one without a deploy, made it possible to try a new behaviour in minutes and see whether the idea held up before anyone built it properly.

How it was built

I built it with coding agents, and the lesson that carried over wasn’t about speed. Left to their defaults, agents quietly narrow the problem. They stub a piece “for now”, defer an edge case, or choose the option that’s quickest to write. The rules that mattered most were the ones against that:

  • never cut or stub scope without my approval
  • pick the best solution first, and treat effort as a feasibility input, not the deciding factor
  • verify how the code currently behaves before proposing a change

Plans were validated before implementation, and code was reviewed before merging. Each feature’s working context lived in a dated folder, so nothing depended on one session’s memory. The same setup later became the template for how I run product management; see the operating model case study.

What carried over

  • Evidence for engineering’s decisions. Questions about routing, state and latency came to the production team with measurements from something people had actually used. The decisions stayed theirs; they just didn’t have to start from a whiteboard.
  • A reference, not a code drop. I wrote up the prototype’s pipeline as a language-agnostic reference for the engineering team: the agents, the routing, and how prompts are assembled. Production was built in its own stack on its own terms, so what transferred was the design and the lessons, not the code.
  • Specs written from evidence. Several product concepts were tried with real users here first, so the requirements that followed described something already shown to work.

What I took from it

  • A prototype is a way of asking a question. Its value is the answer it produces, and the answer outlives the code.
  • Prompt-enforced contracts are fragile wherever exact data matters. Enforce them in code.
  • Real users in real places rearrange your priorities.