Operating model
An AI-native operating model for product
How I keep a pre-launch AI hardware product coherent as its only product manager, across six teams and their partners, 400+ requirements and pilot programmes, while the definition changes weekly. The answer was a system of record built for agents as much as for people, with human judgement designed in on purpose.
The load
This only makes sense at the scale and speed it was built for, so that comes first.
| Surfaces | Glasses hardware and firmware, an AI stack, a phone app and a glasses app |
| People | Six functional teams, plus the partners and stakeholders who depend on what they build |
| Definition | 400+ requirements across hardware, firmware, AI, apps, platform and use cases. More than a quarter were revised within a single quarter |
| Code | About ten production repositories |
| Pace | Pilot programmes secured and both apps in flight, with the definition still moving underneath them |
| Product management | One person |
Why this needed a system at all
On a smaller product, or a slower one, keeping a single current source of truth is simply part of the job, and doing it by hand is fine.
At this scale it breaks in one of two ways:
- The PM becomes a clerk. Documents go stale faster than one person can re-sync them, and product work turns into document maintenance.
- The PM becomes the source of truth. The documents rot, the answers live in one person’s head, and every team queues behind them.
Neither leaves time for the product. This operating model is the third option. Agents do most of the reading, reconciling and checking. People review what they produce. My time goes on the decisions only a person should make.
The starting point
When I started, requirements for the hardware, AI and app layers lived in three different kinds of places: spreadsheets owned by partner teams, tickets in a project tracker, and documents. That caused three problems:
- No single answer to “what does this product need to do?”
- Changes didn’t spread. A change in one layer didn’t reach the layers it affected.
- Dependencies were invisible. You couldn’t trace one user-facing outcome down to everything it depended on.
A closer look at the tracker made it concrete. Edits to a spec notified no one. Requests to partners had no “sent” or “acknowledged” state. Decisions were made in comment threads and then lost there.
The premise: design for two readers
Every artefact has two readers, a person and an agent. Both need to be able to find it, read it and verify it. Anything that needs judgement about the business, the market or users stays with a person, and the system makes that hand-off explicit instead of leaving it to chance.
What’s distinctive
1. Decisions are reviewed like code, and approval is a merge
My first attempt was a hosted database behind a custom tool server for agents. It gave agents access, but it sat outside version control, so nothing in it could be reviewed, diffed, or used to trigger automation.
Our CTO proposed a git-based decisions repository. Before committing, I ran a throwaway proof: could plain files support structured queries and dependency tracing with no server at all? They could. We migrated the whole requirements corpus and cut the agents over in a day.
Records are now markdown files with structured metadata. A small build step generates the only two things git doesn’t provide, a queryable index and a dependency graph. Because it’s just files, any agent can work with it, whatever model is behind it.
The rule that matters most: agents can draft and propose, but never approve. A decision exists only when a person merges the pull request that proposes it, after an automated review. With this many parties, that removes the most corrosive question in an AI-heavy team: “did we agree to this, or did a model write it down?”
2. Knowledge and work live in different places
Knowledge that should last becomes a record. Work to be done becomes an issue. Coordination changes constantly and needs notifications and status; decisions need history and review. One tool that tries to do both does both badly.
So each team keeps its own work repository and pulls its slice of the definition from the shared records, instead of receiving copies. We stopped exporting specs into the project tracker for the same reason: a rendered copy of a spec is exactly the thing that goes stale.
3. Agent output is a claim, not a conclusion
This came from failures, not theory. One review cycle produced overstated findings, understated findings, and one finding that was simply invented. When I checked an unverified follow-up list against the live threads, almost everything on it was already done, already answered, or framed wrongly.
The rules that followed:
- Verify before asserting. “I haven’t verified it” is not the same as “it isn’t there.”
- Scrutinise negative findings harder. A critical review pass is biased toward inventing problems.
- Ground in history, without deferring to it. Significant analysis checks code, decision records and open issues first. It can still conclude a past decision was wrong, but it has to show it knew about it.
Where the cost of being wrong lands on other people, the check is enforced in tooling rather than left to memory. A verification agent checks every follow-up against its live thread before it reaches me. Every file, parameter and route named in an engineering handover is checked against a pinned version of the code before it goes out, so the receiving team never builds against something that has moved.
4. The judgement boundary is designed, not assumed
Most AI workflows leave “when should the agent ask me?” to the model’s instincts. Here it’s routed by what the choice encodes, not how technical it looks:
- Business, market and user judgement comes to me, even when it looks like a methodology detail.
- Execution choices stay with the agent.
- Engineering design belongs to the engineers.
The same boundary governs attribution. Drafts default to one voice, which can make a team decision read like one person’s edict, so artefacts mark whether they’re my proposal or a record of team consensus.
5. A product manager in engineering’s code, as a guest
I write code, and the rules change with the mode. Prototypes I own are built for speed and simplicity. In a repository engineering owns, I’m a guest: I read that repository’s own rules first, work on a branch, review my own change before opening a pull request, and stop for a person on two things, what to do with review findings and whether to merge.
The team’s tooling merges automatically once checks pass. That’s the right default for engineers’ own work. For mine, I deliberately break that chain.
What it made possible
- A definition that kept up. When launch scope was restructured by surface, more than a hundred requirements changed within a quarter. Each change went through review, and every team read the new version from the same place.
- Launch readiness checked against reality. I traced each launch requirement to the code that implements it across the product’s repositories. That turned “are we on track?” into a list of specific gaps that engineering could act on.
- Architecture and requirements reconciled both ways. Where the architecture documents and the requirements disagreed, each conflict was resolved and recorded, not just noted.
- Handovers that don’t waste the receiving team’s time. Verified handovers matter most for the people furthest from the source, who can’t ask a quick question across the desk.
How it evolved
Each change came from the load growing, not from tinkering.
| Change | Why |
|---|---|
| Hosted requirements database and custom agent tool server | Agents needed structured access to requirements |
| Split finished deliverables from work in progress | The repository mixed “what we decided” with “how we got there” |
| Moved to git-based decision records; cut agents over in a day | Records needed review, history and automation |
| Retired the project tracker; one work repository per team, pulling from the shared records | Six teams and their partners needed coordination and knowledge to live apart. Pushing updates out to every team was proposed and declined in favour of teams pulling |
| Verification passes on follow-ups and handovers; less delegation to sub-agents | As more people acted on the output, unverified output cost more than it saved |
What’s still unsolved
- Not every specification lives in it. Some partners keep their own. We link to them rather than copying, so drift at that boundary is caught by people reviewing, not by the system.
- Where prose ends and records begin. How much of a long strategy document belongs inside decision records isn’t settled.
- Configuration rots. Agent rules contradict each other and references go stale. It needs periodic audits like any other codebase.
- Every automation has a cost. Reviews on every change, verification agents and parallel research all cost time and money. The discipline is knowing which ones pay for themselves.
A system like this doesn’t make judgement unnecessary. It makes it visible, so you can see exactly where a person had to decide.