Who thinks, who works, who just runs: how do you build a team of AI models?

Using one AI model for every task costs more than it delivers. A model team splits the roles: one thinks, one works, one just runs once the path is drawn clearly enough. Which modell is enough depends on the task and not on habit. Split it cleanly once and the same work comes back more reliably, for a fraction of the cost. The strongest model available to me right now, Fable, takes the role of strategist and thinker for me. It plans, breaks the task down and makes the decisions that need real context. A solid allrounder takes the role of the worker and carries out what the strategist lays out. A very cheap model takes the role of the pure executor, for subtasks that are described clearly enough. In my day to day these are Fable, Sol and Luna Max. The setup checklist from the same session lists the last two under their official names: gpt-5.6 for demanding work, gpt-5.6-luna for narrow, clear and frequent execution. A model is the single AI system behind a tool. An agent is the role on top of it: one clear job, access to tools and data, several steps in a row. This note is about the models. The role stays in place, and you swap the model behind it as soon as a cheaper one is enough for the same job.

The line-up

  • The thinker plans and decides. Typical jobs: settle a backend architecture, read a finished draft as the reviewer.
  • The worker delivers and keeps going. Typical jobs: carry one thing cleanly through hours of work, break the subtasks down far enough for the runner to understand them.
  • The runner executes what has been chewed up for it. Typical jobs: the daily mail review, rework many drafts by one fixed rule.

Why the breakdown of the task decides

The pure executor only understands what has been chewed up clearly enough beforehand. Deciding a backend architecture or solving an open problem, that is not something it can do. That is what the strategist is for. The real trick sits one level below: I do not instruct the cheap model directly. I instruct the worker to break the subtasks down far enough for a simpler model to understand them cleanly. Breaking things down costs little, because all the execution afterwards sits with the cheapest model. How to find that breakdown inside your own process is in Are Codex and Claude Code coding tools?.

You cannot tell it: build this backend architecture. It only executes.

What this changes per month

Two tools at 200 euros each add up to 400 euros a month. For that you get an architect that feels like working with someone who thinks along, and next to it an executor that costs practically nothing. The cheap model is about 25 times cheaper than the strongest model from the same provider, and cheaper again by a wide margin than the strategist. Before this, I used to burn through my 200 euro limit in two days. Since execution sits with the runner, I no longer reach the limit, even while I keep twenty things going at once.

On a 200 euro plan you can do, I believe, 100,000 runs with Luna Max, and you simply do not reach your limit.

During the Deep Dive itself we ran a whole task live on the cheap model, including browser research and reworked drafts. The entire run cost less than one percent of our allowance. Why that model can be this cheap is itself a team result: the worker model helped build the architecture the runner runs on. The models make the models faster and cheaper.

What this means for your own model team

That takes some humility: not running every small thing through the most expensive model just because you can. Once you have found a way for a simpler model to solve a task reliably, you write that way down as a short instruction. From then on the same task runs through the cheap model every time, without needing the strategist again. The strategist is not locked into one role either. Sometimes I use it as the planner, sometimes as the reviewer over a finished result. Where you stand overall is sorted by the Jarvis ladder, and how to describe a task cleanly in the first place is in Should the AI interview you?.

This appetite for a clean operating system for a team is not new for me. In 2018 I wrote down fixed communication rules for a purely volunteer team: short messages in the main channel, a daily searchable summary with a date for each area, and conflicts never in the group chat, only one on one. Today’s model team runs on the same principle. Only the members are no longer people, they are models with clear roles.

On record

KI DeepDive Agentic Mindset, public live session, 3 August 2026.

  • Three roles in the model team: the strongest model as strategist, an allrounder as worker, the cheapest model as pure executor
  • Architecture decisions and open-ended problem solving stay with the strategist
  • The cheap model is about 25 times cheaper than the strongest model from the same provider
  • On a 200 euro plan that allows a very high number of runs without reaching the limit
  • The daily mail review runs entirely on the cheap model
  • The live run during the Deep Dive cost less than one percent of the allowance

Model team line-up

Who thinks, who works, who just runs: the three roles as a template to fill in. PDF · A4, 1 pages, 31 KB, as of 3 September 2026.

This note keeps growing

2026-09-03: Deepened with my own team communication rules from 2018: the same operating-system thinking, back then for a volunteer team, today for a model team.

2026-09-03: Deepened: the line-up with two typical jobs each, the breakdown routed through the worker instead of straight to the runner, price ratio and the 200 euro plan with quotes, the live run from the session, source list.

2026-09-02: Planted from the Agentic Mindset Deep Dive.

ZukunftBilden GmbH · Salzburg · +43 681 81655313 · office@zukunftbilden.eu