BLOG


Cover — in the same restrained, icon-based systems style as “Connecting the Dots,” a dispatch board routes different jobs among a person, a Python script, a small local model and a frontier model. Small gauges show that cost is measured in time, memory, quota and attention—not tokens alone.
Editorial systems illustration of tasks being allocated among human, deterministic, local and frontier intelligence.

The question is no longer whether intelligence is available, but where it is worth spending.

The Bit & Harness

Agentic Production and Resource Allocation

Posted on September 30, 2026 by Peter Loomis


Introduction

Last month, I wrote about some of my recent experiments integrating ChatGPT, Codex, local models and Python into a more connected development process. What started as a few simple automations had expanded into something much larger, with different tools and models doing different parts of the work. By the end, I was already thinking more carefully about what should remain local, what needed a frontier model and what could simply be handled deterministically without AI at all.

Well, I’ve kept going. Over the past month, I’ve been running more of these things at the same time, building out some of the ideas and putting them through actual use. Along the way, the question has shifted again. I am less interested now in whether AI can do a particular thing. Usually it can, at least in some form. I’m becoming much more interested in which intelligence should do it, what that intelligence actually costs and whether using it there is worth what comes back.

Once intelligence becomes heterogeneous, the problem starts shifting from obtaining intelligence to allocating it.


2. Running the Farm — a three-panel icon illustration using the prior post’s visual language: a Stampede sends many agents forward at once, a Raid uses a smaller targeted team and a Buggy Parade jams the route with too many workers competing for the same machine.
A desk inside an urban apartment with workstations open and server rack to the right, the city visible through a window in the background.

Parallel intelligence can look like a team, a tactical group or a traffic jam.

1. Picking Up Where I Left Off

At the end of last month, I had arrived at a hybrid process almost by accident. Having burned through my Codex quota, I moved parts of the development back into ChatGPT, started experimenting more seriously with local models via Bionic and increasingly relied on Python for things that did not need a model at all.

At the time, I was mostly trying to keep the work moving. Codex was extremely capable but limited by quota. ChatGPT was useful for thinking through architecture and figuring out what to do next. Local models offered another source of intelligence without watching a credit balance disappear every time I asked them to do something. Python could handle predictable processing cheaply and repeatedly. However, I was still carrying files and instructions around myself.

Different roles started emerging. I started using Codex more deliberately for higher-value implementation work. ChatGPT became increasingly useful for discussion, architecture and figuring out how larger pieces should fit together. I experimented with several local models, including different sizes for different kinds of work and started running batches of multiple local agentic jobs at the same time.

Somewhere in there I realized I was no longer just choosing between AI products. I was allocating and routing work between different kinds of smart resources.


2. Running the Farm — A stampede of horses filling the frame running from left to right.
A stampede of horses filling the frame running from left to right.

Parallel intelligence can look like a team, a tactical group or a traffic jam.

2. Running the Farm

There is something incredible about having several AI agents working at once. One process can be examining a problem while another is implementing something and another is chewing through a long, relatively inexpensive local job. Meanwhile I can be talking through the architecture somewhere else, reviewing output from something that finished earlier or just doing something myself.

In the process, I started developing my own language for some of this. For example, releasing a bunch of local agents at once has become a 'stampede.' A smaller targeted group has become a 'raid.' When I get carried away and release too many that the machine starts slowing down, that is a 'buggy parade.'

The terminology is funny, but the problem underneath it is real. Initially, local models felt abundant. Since there wasn’t a quota or token meter running in the corner of the screen. I could assign several things, let them run for hours and even go to bed without worrying about what those hours were doing to my weekly allowance. Then I woke up and saw the machine still churning.


3. The Intelligence Was Free — a close, simplified Activity Monitor–style composition: four processes remain active while a 64 GB memory gauge is filled past 60 GB, swap and compression indicators rise and ordinary workstation apps wait outside the frame. This could also be based on a real screenshot if one is available.
Black & white editorial photograph of children on a seesaw in a park.

The intelligence was free. The workstation wasn’t.

3. The Intelligence Was Free

One morning I had four local processes still working after running through the night. My Mac has a finite of memory and was using more than 95% of it. Several more gigabytes had moved into swap. Memory was compressed. The processes were still chugging along. I hadn’t even reopened my normal Firefox session yet. And this is my workstation, where I also need to design, write, browse, communicate, edit images, listen to music and do everything else.

A few days earlier I had been running a similarly heavy collection of processes and later had a kernel panic after trying to resume a YouTube mix I’d been listening to. I can’t say the AI workload caused the crash, but it certainly made me more conscious of what I was asking one machine to carry. Local inference may not have been costing me additional tokens, but it was consuming almost everything else. The intelligence was "free," but the workstation wasn’t.

That distinction has become increasingly important. Local models consume memory, storage, compute and time. They can also occupy my computer which I need for lots of other things. A model that spends ten hours processing something while I sleep may be a fantastic use of otherwise idle capacity. The same model still running at noon while I am trying to work has a completely different cost. So what does “cheap” actually mean?


4. Cheap Depends on What You Need — two routes to the same finished result. One uses a scarce frontier model for twenty minutes; the other uses a local model for eight unattended hours. Surround each route with small icons for quota, elapsed time, supervision, hardware load and rework instead of a single dollar price.
Editorial photo of a darkened office building with lights and computers still running.

The least expensive worker depends on what else the work consumes.

4. Cheap Depends on What You Need

Last month I was already thinking about preserving frontier models for work that actually required their capability. Now, the calculation has become more nuanced.

A local model that runs for eight hours and produces something I have to substantially redo may be more expensive in practice than spending twenty minutes of a scarce premium resource to get a better result. At the same time, using high-quality coding capacity to grind through something repetitive just because it can do it quickly can be wasteful if the same job could run locally overnight without affecting anything else.

So, the amount of time something takes also means something different when I’m not waiting for it. If a frontier model can finish something in eight minutes and a local model takes forty-five, the frontier model is obviously faster. But if I’m asleep, eating dinner or working on something completely different, those extra thirty-seven minutes may cost me almost nothing. In that case, a slower resource can sometimes be the more economical one simply because the work is happening asynchronously.

This has made me re-evaluate what I really want from all of this. The fastest possible answer is not always the goal. There is something appealing about a quiet stream of useful computation running underneath everything else I’m doing, advancing work while my attention is somewhere else. Of course, if enough of those streams get run simultaneously, I'll pay for it with a buggy parade.

There are even times when I am the cheapest available intelligence. I have spent enough of my career doing production work to know that sometimes it is faster to just do the thing myself. If I can manually clear a bottleneck in twenty minutes while another process is occupied, I am not going to spend two hours engineering an elegant automation simply because automation sounds more sophisticated. I may automate it later if the task repeats enough to justify the effort.

This is where token pricing stops being a particularly useful way to understand the economics. The actual cost includes capability, latency, memory, compute, quota, reliability, supervision, human attention and the opportunity cost of whatever else could be happening instead.

Even availability is contextual. A local model may technically fit into memory, but that does not necessarily mean starting another one is a good idea. A frontier coding agent may technically be available, but if I am approaching a quota limit I may want that capacity for something more important later. So, available doesn’t always mean economical.


5. More Workers, More Management — several agent windows finish around one person who is physically carrying result packets between them. The workers are idle while the central human communications layer sorts, remembers and reroutes each handoff.
Editorial photo of five cats on the grass.

Intelligence can become abundant faster than coordination does.

5. More Workers, More Management

There is another cost I didn’t fully appreciate when I first started experimenting with parallel agents. Someone has to run the farm.

Last month I wrote about becoming the bottleneck between different threads, copying information around, authorizing actions and trying to keep several streams of development moving. I was already trying to reduce that because it was maddening. Over the last month the problem has become easier to see because there is simply more happening.

A worker finishes something. I notice that it finished. I find the result. I remember what larger piece of work it belongs to. Maybe I bring that result into another conversation to figure out what it means. Then I take whatever comes out of that discussion and carry it somewhere else so the next worker can continue.

That can work surprisingly well. But, it turns me into a bottleneck as the communications layer.

As the number of workers increases, the coordination around them increases, too. More parallel intelligence can create more work for trying to direct it. If I save two hours of production but spend an hour trafficking context between five different places, some of the supposed gain has quietly disappeared.

To address this, I’ve been working on ways to make context, work state and results more persistent so that individual sessions don’t have to carry everything themselves. Although I’m still figuring out what that system should become, the pressure behind it is increasingly obvious. Intelligence can become abundant faster than coordination. And, consequentially, the cost of intelligence also includes the cost of coordination.

"There is nothing so useless as doing efficiently that which should not be done at all."

— Peter Drucker

6. Who Gets the Job? — a routing junction sorts task cards toward a human, Python, a small local model or a frontier model. Each route rejoins at a testing gate; failed work loops back for another attempt or escalates to a more capable resource.
Editorial image of an auction in a modern office building with people holding up signs with frontier LLM logos.

The intelligence that does the work does not have to be the intelligence that judges it.

6. Who Gets the Job?

At first, most of these allocation decisions were intuitive. If it looked complicated, I'd give it to Codex. More architectural? I'd talk it through with ChatGPT. Repetitive and ok to run overnight? Give it to a local model. Doesn’t need AI at all? Let's write a script. If it will take longer to explain than to fix, I'll just do it myself.

I am, however, becoming expecially interested in what happens when those decisions are informed by actual evidence rather than intuition alone. If one category of work consistently runs successfully on a smaller local model, that becomes useful to know. If another repeatedly fails locally but completes quickly with a more capable model, that is useful too. How long did it run? How much memory did it occupy? Did it require intervention? Did the result survive review? Did I have to run it again somewhere else?

There is another distinction hiding in that last question. The agent that performs the work does not necessarily have to be the judge of whether the work is correct.

A less expensive model might be perfectly capable of attempting a bounded piece of work, particularly if the result can be tested afterward. If it passes, great. If it fails, maybe it gets another attempt, moves to something more capable or eventually comes back to me. That changes the calculation considerably. I don’t necessarily need the smartest available intelligence to perform every task if I have a reliable way of determining whether the result is acceptable.

In some cases that determination doesn’t require intelligence at all. Software can test whether something exists, whether it runs, whether a value falls inside an expected range or whether an output matches known criteria. Other results are ambiguous enough that another model or a human needs to look at them. That begins to separate the cost of doing the work from the cost of knowing whether the work was done well. Over time, those experiences start describing the actual economics of the work.

The same is true of machine state. Having enough free RAM to technically launch another worker doesn’t tell me whether I should. Swap, memory compression, whatever else is running and what I plan to do next all matter. A computer that can launch another model may still be better off waiting.

It feels less like picking a favorite AI and more like directing production workflow considering that different jobs require different capabilities and the best worker depends partly on the conditions around the work.


7. The Human Cost — a two-layer diagram separates an intelligence deciding what it wants to do from an authority gate deciding what it may do. Beside it, several independent work branches continue while only one pauses at a small “Peter needed” decision point.
Editorial photo of a man and woman looking at the computer screen.

Often I start work before doing yoga, which I continue to monitor throughout the day. Then, I do not finish yoga until late in the day.

7. The Human Cost

As these experiments have become more agentic, another constraint has become increasingly obvious: my attention. Any worker is not particularly inexpensive if I have to watch it constantly. A process that repeatedly stops for routine permission or needs me to interpret every intermediate result may consume very little computationally while remaining relatively expensive in the one resource I have the least ability to scale.

There is also a difference between giving something intelligence and giving it authority. I’ve had local agents request access to places I had absolutely no intention of letting them go. Nothing happened because I was there to stop it, but it was a useful reminder that an AI deciding what it wants to do and a system deciding what it is allowed to do are different problems.

Increasingly, I'm interested in keeping those two things separate. An agent can reason about a problem without automatically receiving unlimited authority to act on whatever conclusion it reaches. Some decisions are routine. Some genuinely need me. Figuring out the difference is part of reducing the human cost while still remaining at the top of the loop.

So, I’ve started thinking about my own attention like another dependency in the work. If a branch reaches something that genuinely needs my judgment, that branch may have to wait for me. It doesn’t necessarily follow that everything else should stop too. Other work that doesn’t depend on that decision should continue.

The distinction may sound small, but it changes the relationship considerably. “This needs me” is different from “the whole system stalls without me.” Ideally I can walk away for a while and come back to the relatively small number of decisions that actually required my involvement while other work continued without me. It's another kind of allocation to be made. Not only which model gets the work, but also which decisions truly need my attention.


8. Treading Water — a deliberately simple comparison: four agents × eight hours produces a large “32 agent-hours” activity counter, while a much smaller tray labeled “useful results” waits to be evaluated. The visual question is whether the activity moved the marker forward.
Editorial black and white image of dark water with waves disrupting the surface.

When activity becomes cheap, progress becomes the more important measure.

8. Treading Water

There is a slightly uncomfortable side to all of this. Some days there is an incredible amount happening on my computer. Multiple agents are running. Code is being written. Files are changing. Interfaces are appearing. Ideas that might once have stayed in a notebook are turning into working prototypes. Sometimes I can literally go to sleep while several things continue developing. It feels productive. But is it?

I’ve found myself looking at all this activity and still wondering where I am. Am I actually moving something forward or just becoming extremely efficient at generating more things to manage? Am I building something useful? Am I learning? Am I making something commercially viable? Is this development opening possibilities I couldn’t reach before or am I just treading water considerably faster?

I don’t think those questions invalidate the work. Some of the tools I’ve been developing have already done useful things for me in the real world, and the amount I’ve learned by pushing them into actual use is substantial. But the distinction between activity and progress matters even more when activity becomes cheap.

Four agents working for eight hours could represent thirty-two agent-hours of activity inside a single night’s elapsed time. That sounds impressive until you look at what came back. Were those hours worth it? That may be the more meaningful economic question.


9. Allocating Intelligence
Editorial black and white image of dark water with waves disrupting the surface.

When activity becomes cheap, progress becomes the more important measure.

9. Allocating Intelligence

For most of my career I have been deeply involved in production. I have designed the thing, built the thing, resized it, exported it, renamed it, uploaded it, QA’d it and fixed whatever broke. Working with developers and production teams taught me how to direct larger systems, but I have always remained comfortable jumping back into the work myself.

AI is changing the scale of that relationship. I can now have several different forms of intelligence working simultaneously, each with a different combination of capability, speed, expense, persistence and access. Some of them can continue after I stop working. Some are vastly more capable than others. Some are almost free to run but expensive in time or hardware. Some are scarce enough that I have to think carefully about where I spend them.

My role moves around inside that environment. Sometimes I’m designing the architecture. Sometimes I’m directing. Sometimes I’m reviewing. Sometimes I’m routing information. Sometimes I’m still the production artist doing the stupid repetitive thing because, at that particular moment, I am genuinely the most economical worker available.

What is changing is that those decisions are becoming conscious. The question is no longer simply whether I can automate something, or whether an AI can perform the task. I am starting to ask what resource should do what, under what conditions and at what cost. Eventually some of those allocation decisions may themselves become more automated, informed by what actually happened the last ten times a similar job was run.

And increasingly, I’m interested in what can continue without me. Not because the goal is to remove myself from the process, but because my attention is probably better spent deciding what deserves to happen, evaluating consequential results and changing direction when necessary than repeatedly moving routine work from one step to the next.

But there is a danger there too. Keeping every worker busy would be easy enough. That doesn’t mean they should all be working. Utilization isn’t necessarily value.


Conclusion
Editorial black and white image of dark water with waves disrupting the surface.

The harder question is becoming where to spend it.

Conclusion

Last month I ended up with a hybrid process because the limitations of each individual platform kept forcing me somewhere else. A month later, that compromise is starting to look more like the point.

There may not be one ideal intelligence for the work. There may be a shifting collection of them, local and remote, large and small, deterministic and inferential, automated and human. The interesting part is increasingly how they are combined, when they are used and whether the work they produce justifies the resources they consume.

I’m still building and still pushing this through actual use, which means I’m sure the economics will keep changing as the tools do. But for now, I have more intelligence available to me than I’ve ever had before. The challenge is becoming how to leverage it all for high value tasks as continuously, automating as much as possible, including resource allocatiion and routing the right models on the right things and giving me a way to monitor everything as it passes through the pipeline.

Is that so much to ask? Stay tuned!



You may enjoy some other AI, creativity and systems-related posts from our blog:


New Project?

Planning a new creative project and need expert support? We can help. Contact us today!


Stay Connected

Follow us on Facebook, Instagram, LinkedIn, X and YouTube for more design content and inspiration! Subscribe to our mailing list for our email newsletter with periodic updates