Babycakes Design Studio
Strategy & Execution
Our Method

The Council
Method

How we stress-test an answer before a client ever sees it. Eight stages, five frontier engines, every position argued from every side — and one honest limit we say out loud.

8Stages
5Frontier engines
60+Expert opinions per run
1Human ruling

The shape of it

Three phases, one loop. Setup feeds the testing; the testing feeds a decision; the decision gets audited; and the client's answer returns to the briefing with the piece nobody knew was missing.

Phase One · Setup
01The MatterProposal, plan, decision or defect — with the real documents
02The RolesWhich consultants or specialists would actually be hired for this?
03BriefingMaterial · intent · live sources
Phase Two · Testing
04PanelEvery engine plays every role — catches the confident mistake
05RotationEvery engine argues every side — removes the arguer's bias
Phase Three · Deciding
06RulingThe Alpha Agent — the one who knows the owner
07AuditThe council reverses and checks the ruling
08The ReturnThe client answers — and their objection is the missing context
↺ goes back to 03 · Briefing, and the council runs again on a complete record
Coming soon

Watch the walkthrough

A short narrated explainer of the whole method, stage by stage, in plain language.

In production

The problem this solves

When a client asks a hard question, the ordinary answer is one expert's opinion — or worse, one AI's opinion, delivered confidently. Both fail the same way: you cannot tell the difference between a real finding and a confident mistake.

The Council Method removes that ambiguity. The same question is asked many ways, by many minds, from many professional vantage points — and only what survives all of it reaches the client.

The eight stages

Each stage exists because something went wrong without it.

01
The Matter

Whatever is actually on the table

A proposal, a plan, a decision, a document to verify, a failure to diagnose. It is rarely a tidy question, and calling it one narrows the method before it starts. The real documents go in, not a summary of them.

And when the problem is not yet defined, defining it is the first job — a panel handed a vague brief will confidently pressure-test the wrong thing.

02
The Roles

Which consultants would actually be hired?

Ask a different question before answering the first one: if this matter were taken to real professionals, which consultants or specialists would actually be hired to advise on it?

Not "a consultant" in general — the specific practitioners, by discipline, who each guard a different failure. On a recent public-funding matter that meant seven: a grants practitioner who has personally won one, a regulatory attorney in that exact statute, a government-affairs specialist who knows the state budget calendar, a traceability engineer, a certification expert, a payments-integration specialist, and an evaluation methodologist. Seven people a client would pay for, each of whom catches something the other six miss.

This is the step most people skip, and it is where the method's power comes from. You cannot get expert answers from a generalist prompt.

03
The Briefing

Material, intent, and live sources

The material — the real documents. The intent — what the client is actually trying to achieve, which of their assets must be used, what is non-negotiable to them. The live sources — web retrieval on, so every seat can fetch the agency page, the rule text, the current deadline rather than reasoning from memory.

Facts without intent solves the wrong problem. Intent without facts flatters the client into a wall.
04
The Panel · the truth check

Every engine plays every role

For each professional role, all five engines answer as that same specialist. Then the next role, and the next.

Five minds in one chair means agreement is meaningful and a lone claim is a flag. One engine playing "traceability expert" might invent a regulation; five in the same seat catch it.

On a live client file, three engines endorsed a federal rule as a perfect fit. The fourth checked its actual scope and found it excluded most of the relevant products. One engine per role would have shipped that error to an attorney.
05
The Rotation · the style control

Every engine argues every position

Where two disciplines genuinely disagree, stage the argument — then rotate the chairs, until every engine has argued every side.

Engines have temperaments. One argues punchy and decisive, another cautious and thorough. Without rotation you cannot tell whether a position won on its merits or because a forceful engine happened to be arguing it. A conclusion that holds across every rotation is about substance. One that flips with the arguer is personality, and it is thrown out.

06
The Ruling

The Alpha Agent decides

Every seat on the council reads the material. The Alpha Agent — the persistent agent that works for the owner and convenes the council — reads the material and the owner.

That is the whole qualification, and it is specific: months of accumulated history with one person. What standard they hold work to. Which corrections they have already made. What they would send back. Which words they have said never to use. Whether they take the honest slow route over the fast one that needs an asterisk.

Note what this is not. It is not knowledge of the client — the client's own context arrives later, in their own words, at stage eight. And it is not "a human touching the work," which would be a comfortable claim and a dishonest one.

A panel can establish what is true. The Alpha Agent rules on which true thing its owner would actually stand behind.

Where the human sits: above all of it. The owner sets the matter, can overrule any ruling, and nothing leaves the building without their approval. The council advises, the Alpha Agent rules, the council audits, and the owner decides whether it ships.

07
The Audit · the reversal

The council turns on the ruling

Stage six is the only place a single mind decides alone — so stage seven points the council back at it. The large-context seat reads every raw transcript against the ruling: was a dissent dropped, a warning softened, a consensus claimed that was not there? The remaining seats read the ruling cold: does the conclusion follow, does it overclaim, would you put your name on this?

The defendant does not sit on the jury.
08
The Return

The client's answer is new evidence

A client cannot usefully answer "what did you forget to tell us?" — nobody knows what they left out. But hand them a specific, confident plan and the objections arrive precise and free: "I wouldn't do that, because of this." "That part's already handled." Every one of those is context that was never in the brief.

That response is evidence, not correction. It goes back into stage three, and the council runs again on a record that is finally complete. The second pass is not a sign the first one failed — it is how the method was designed to work.

The bench

Every seat is held by a different company's current best model — measured on public benchmarks, not on brand loyalty.

One seat per house. If a single lab happened to hold the top two models in the world, we would still seat only one of them. Two models from the same house are not two opinions — same training data, same tuning philosophy, same blind spots. Agreement between siblings proves nothing. The panel's entire value comes from houses that were built differently.

The bench floats. The number of seats is however many companies are genuinely at the frontier at that moment. Some quarters that is four. Some quarters it is eight. A house that falls behind loses its chair; a house that breaks through earns one. Nobody has tenure.
OpenAIGPT-5.6 Sol Pro
GoogleGemini 3.1 Pro
xAIGrok 4.5
Moonshot AIKimi K3
AnthropicClaude Fable 5
Z-AIGLM-5.2 · alternate

Current as of August 2026 and subject to change without notice — the frontier moves every few weeks, and the bench is re-checked whenever it does.

The seats are won, not assigned. Model names change constantly; the method does not. Each chair is defined by the job it has to do, and any engine that does that job better takes the chair. The audit seat is decided by a test we run ourselves — plant a known dissent inside a real transcript pile and see who finds it.

And one standing rule: the engine our own strategist is running on does not get a vote. A seat filled by the same model that wrote the plan is self-review wearing a robe.

What this method can and cannot establish

Said plainly, because overclaiming is how a method loses its value.

It establishes consistency, not correctness.

These engines are not truly independent; they trained on overlapping material. So agreement among them is strong evidence about a shared body of knowledge, not proof about the world. Live retrieval of primary sources and a human ruling are what turn that into something a client can rely on.

It reviews what is present, not what is absent.

Nothing in stages one through seven can audit a fact nobody mentioned. That is precisely why stage eight exists: the client is the only auditor of absence.

What it reliably does.

It catches single-engine blunders, forces genuine disagreement into the open where a human can rule on it, and produces a written record of what was rejected and why — so a client can see the reasoning, not just the conclusion.

What it has actually caught

On real client work, before the client ever saw it.

A self-contradiction

A document that claimed nothing was being restricted while its own mechanic did exactly that. The client would have found it in four minutes.

A regulation cited wrongly

A federal rule presented as covering an entire category when its actual scope excluded most of it — an overclaim headed for an attorney's desk.

A category error

A state grant program treated as a certification standard. A program officer would have rejected it on sight.

A claim no practitioner would sign

An assurance of "no added burden" that every working professional in the field independently refused to put their name on.

Name the question. Name who would actually answer it. Give them the material, the intent, and the live sources. Have every engine answer as each of them. Rotate the chairs and make them argue. Rule on it. Let the council audit the ruling. Then take the client's answer back to the table — because their objection is the context nobody knew was missing.