Introduction
Last month, I wrote about some of my recent experiments integrating ChatGPT, Codex, local models and Python into a more connected development process. What started as a few simple automations had expanded into something much larger, with different tools and models doing different parts of the work. By the end, I was already thinking more carefully about what should remain local, what needed a frontier model and what could simply be handled deterministically without AI at all.
Well, I’ve kept going. Over the past month, I’ve been running more of these things at the same time, building out some of the ideas and putting them through actual use. Along the way, the question has shifted again. I am less interested now in whether AI can do a particular thing. Usually it can, at least in some form. I’m becoming much more interested in which intelligence should do it, what that intelligence actually costs and whether using it there is worth what comes back.
Once intelligence becomes heterogeneous, the problem starts shifting from obtaining intelligence to allocating it.
Parallel intelligence can look like a team, a tactical group or a traffic jam.
1. Picking Up Where I Left Off
At the end of last month, I had arrived at a hybrid process almost by accident. Having burned through my Codex quota, I moved parts of the development back into ChatGPT, started experimenting more seriously with local models via Bionic and increasingly relied on Python for things that did not need a model at all.
At the time, I was mostly trying to keep the work moving. Codex was extremely capable but limited by quota. ChatGPT was useful for thinking through architecture and figuring out what to do next. Local models offered another source of intelligence without watching a credit balance disappear every time I asked them to do something. Python could handle predictable processing cheaply and repeatedly. However, I was still carrying files and instructions around myself.
Different roles started emerging. I started using Codex more deliberately for higher-value implementation work. ChatGPT became increasingly useful for discussion, architecture and figuring out how larger pieces should fit together. I experimented with several local models, including different sizes for different kinds of work and started running batches of multiple local agentic jobs at the same time.
Somewhere in there I realized I was no longer just choosing between AI products. I was allocating and routing work between different kinds of smart resources.
Parallel intelligence can look like a team, a tactical group or a traffic jam.
2. Running the Farm
There is something incredible about having several AI agents working at once. One process can be examining a problem while another is implementing something and another is chewing through a long, relatively inexpensive local job. Meanwhile I can be talking through the architecture somewhere else, reviewing output from something that finished earlier or just doing something myself.
In the process, I started developing my own language for some of this. For example, releasing a bunch of local agents at once has become a 'stampede.' A smaller targeted group has become a 'raid.' When I get carried away and release too many that the machine starts slowing down, that is a 'buggy parade.'
The terminology is funny, but the problem underneath it is real. Initially, local models felt abundant. Since there wasn’t a quota or token meter running in the corner of the screen. I could assign several things, let them run for hours and even go to bed without worrying about what those hours were doing to my weekly allowance. Then I woke up and saw the machine still churning.
The intelligence was free. The workstation wasn’t.
3. The Intelligence Was Free
One morning I had four local processes still working after running through the night. My Mac has a finite of memory and was using more than 95% of it. Several more gigabytes had moved into swap. Memory was compressed. The processes were still chugging along. I hadn’t even reopened my normal Firefox session yet. And this is my workstation, where I also need to design, write, browse, communicate, edit images, listen to music and do everything else.
A few days earlier I had been running a similarly heavy collection of processes and later had a kernel panic after trying to resume a YouTube mix I’d been listening to. I can’t say the AI workload caused the crash, but it certainly made me more conscious of what I was asking one machine to carry. Local inference may not have been costing me additional tokens, but it was consuming almost everything else. The intelligence was "free," but the workstation wasn’t.
That distinction has become increasingly important. Local models consume memory, storage, compute and time. They can also occupy my computer which I need for lots of other things. A model that spends ten hours processing something while I sleep may be a fantastic use of otherwise idle capacity. The same model still running at noon while I am trying to work has a completely different cost. So what does “cheap” actually mean?
The least expensive worker depends on what else the work consumes.
4. Cheap Depends on What You Need
Last month I was already thinking about preserving frontier models for work that actually required their capability. Now, the calculation has become more nuanced.
A local model that runs for eight hours and produces something I have to substantially redo may be more expensive in practice than spending twenty minutes of a scarce premium resource to get a better result. At the same time, using high-quality coding capacity to grind through something repetitive just because it can do it quickly can be wasteful if the same job could run locally overnight without affecting anything else.
So, the amount of time something takes also means something different when I’m not waiting for it. If a frontier model can finish something in eight minutes and a local model takes forty-five, the frontier model is obviously faster. But if I’m asleep, eating dinner or working on something completely different, those extra thirty-seven minutes may cost me almost nothing. In that case, a slower resource can sometimes be the more economical one simply because the work is happening asynchronously.
This has made me re-evaluate what I really want from all of this. The fastest possible answer is not always the goal. There is something appealing about a quiet stream of useful computation running underneath everything else I’m doing, advancing work while my attention is somewhere else. Of course, if enough of those streams get run simultaneously, I'll pay for it with a buggy parade.
There are even times when I am the cheapest available intelligence. I have spent enough of my career doing production work to know that sometimes it is faster to just do the thing myself. If I can manually clear a bottleneck in twenty minutes while another process is occupied, I am not going to spend two hours engineering an elegant automation simply because automation sounds more sophisticated. I may automate it later if the task repeats enough to justify the effort.
This is where token pricing stops being a particularly useful way to understand the economics. The actual cost includes capability, latency, memory, compute, quota, reliability, supervision, human attention and the opportunity cost of whatever else could be happening instead.
Even availability is contextual. A local model may technically fit into memory, but that does not necessarily mean starting another one is a good idea. A frontier coding agent may technically be available, but if I am approaching a quota limit I may want that capacity for something more important later. So, available doesn’t always mean economical.
Intelligence can become abundant faster than coordination does.
5. More Workers, More Management
There is another cost I didn’t fully appreciate when I first started experimenting with parallel agents. Someone has to run the farm.
Last month I wrote about becoming the bottleneck between different threads, copying information around, authorizing actions and trying to keep several streams of development moving. I was already trying to reduce that because it was maddening. Over the last month the problem has become easier to see because there is simply more happening.
A worker finishes something. I notice that it finished. I find the result. I remember what larger piece of work it belongs to. Maybe I bring that result into another conversation to figure out what it means. Then I take whatever comes out of that discussion and carry it somewhere else so the next worker can continue.
That can work surprisingly well. But, it turns me into a bottleneck as the communications layer.
As the number of workers increases, the coordination around them increases, too. More parallel intelligence can create more work for trying to direct it. If I save two hours of production but spend an hour trafficking context between five different places, some of the supposed gain has quietly disappeared.
To address this, I’ve been working on ways to make context, work state and results more persistent so that individual sessions don’t have to carry everything themselves. Although I’m still figuring out what that system should become, the pressure behind it is increasingly obvious. Intelligence can become abundant faster than coordination. And, consequentially, the cost of intelligence also includes the cost of coordination.




