KT Gilcrease
AI Orchestration · Case Study
Full case study

AI Planning Room

AI Orchestration · Case Study
← Back to work

The idea

Running more than one AI model on a project gets messy fast. Each one lives in its own tab with its own version of the context, nobody is tracking what it all costs, and nothing stops an agent from charging ahead on work nobody approved. I wanted the opposite: one room, one shared transcript, three models from three different companies — Codex, Claude, and Gemini — working as a team, and one person clearly in charge.

The team

None of the models is locked into a job. I assign each one a role — engineer, tester, architect — and those roles come with real job descriptions, the same kind you'd write for a person. I adapted them for agentic models, but most of the expectations and work processes carried over almost unchanged. A job description helped each model understand its lane, and a clear definition of done turned out to be one of the most useful things in the whole setup. The roles are still mine to hand out: if I want Codex on architecture review instead of Claude, it does. They don't get attached to a title; they wear the hat I give them.

One rule holds above everything else: I'm the only one who gives marching orders. No model can tell another model what to do. They can propose, recommend, suggest, and debate — but they only act on direction from the human in the loop.

Debate over agreement

The part I enjoyed most was watching the models debate — the best approach for right now versus the right long-term strategy. They had no problem calling each other out, and they were just as good at coming back together on the same page. That debate became a guardrail of its own. Models tend to over-agree with the human in the loop: a person states a view, and the model lines up behind it. When the pushback came from another model instead of from me, there was far less of that. So when I sensed that risk, I held my own opinion back and threw the question to the team to keep working until they agreed — then bring me one to three options with pros and cons. The final call was always mine, made from options that had already survived an argument.

Why three companies, not one

Yes, you can run multiple agents off a single model — but they all share the same tendencies and the same blind spots. You can't put all your eggs in one basket. In my experience, Gemini handles large contexts and long documents better than the others; OpenAI's models are strong on concepts and work like a senior developer; Claude is strongest at the architect level; and Gemini and Codex both do well with UI and take direction really well. Having all three at once means more options, better short-term versus long-term thinking, and more control over cost.

That shows up in the spending. Claude costs more, so I leaned on it early and for anything risky, security-heavy, or architecture that would carry a lot of features forward. Once the architecture was solid, Claude took a back seat and the work moved to the cheaper models — with full confidence, because the specs, design patterns, and automated tests left very little room to drift. I'd ask for the minimum changes required, with proper error handling, brought back to the room for review. I seldom had to say, "You only implemented a fourth of what I gave you." It was a very different experience, and it didn't feel like a waste of my time.

How we worked

Agile was the heart of it, but the room let us do what most agile teams never have time for: the heavy work up front. For each project, the models and I worked through the methodology, then fleshed out requirements, use cases, specs, and architecture together. We turned those into work orders, and once all four of us agreed a work order was ready, we decided who would build it. Larger work orders were split into phases with deliberate stop points — review, test, and iterate until we all agreed it was ready to move on.

Work ran in parallel when it safely could. One or two models might go off to build while the rest of us planned the next phase; the deciding question was whether two features could be worked on at once without stepping on each other's toes. When the builders came back, a new spec wasn't simply handed to them. We asked for their concerns and questions, whether they agreed with the approach, and whether they knew a better way — and only then finalized it.

Code review was peer review, model to model, and its main job was scope: each model did exactly what the work order said, and nothing more. When a model wasn't catching on, I had another step in and take the reins, then decided whether to hand the work back or let the new model finish the change set.

Shipping ran like a real engineering team with continuous integration and deployment. Nothing auto-committed. We reviewed every change together before committing and promoting it, merged deliberately, and built a deployment package. Automated tests ran before and after deployment, and I did a final run-through as the last user acceptance tester — a sanity check that nothing was missed and nothing needed rolling back.

How I built it

It's a single Node.js runtime — command line first, with a minimal web view that only runs locally — talking to the OpenAI, Anthropic, and Google APIs through separate adapters. State and transcripts persist on disk, and project presets let the room switch between codebases. Control was the design center: only I can issue commands, and anything a model says is treated as plain text, never as an instruction. Every call is costed per model, and a model that crosses its budget is muted automatically, with the reason written into the transcript.

Some of that control I learned the hard way. AI models are built to be over-helpful, and my early setup gave them more room than I realized — they replied so fast, and so repeatedly, that I couldn't catch things quickly enough. Tightening permissions turned out to matter as much as adding features.

My role

Concept, control-plane design, architecture, the cost-governance model, and the team workflow itself — job descriptions, methodology, work orders, phase gates, and final user acceptance testing.

What's next

I'll admit it: this is one of the most interesting things I've built, and one of the most fun — and I still work this way today. The Planning Room is deliberately local and single-owner, and it's meant to stay that way; it isn't a hosted service and won't become one. What it can be is a pattern — something I could help an organization set up inside its own environment. Closer in, the transcript still needs a retention strategy.