A cluttered desk setup with several AI coding assistant windows open side by side, each mid-session, next to an unread pull request and a design document Image credits: Google Gemini
Engineering and Development

Nobody Trained Me to Manage Five AI Agents at Once

Five things are open on my screen and none of them are code I’m writing.

Two separate AI coding sessions running against two different projects, each a few turns into a task I handed off an hour ago.

More than a couple of pull request sitting in my review queue that I’ve opened three times without actually finishing and a design doc a colleague sent this morning that I’ve skimmed but not really read.

And a potential bug I found while I was experimenting with something that I still haven’t reproduced, because every time I go looking for it, one of the sessions pings me back with something to check.. 🙃

I haven’t hand-written more than twenty lines of code today. If a certified vibe coder reads this they’ll probably come at me with “writing code by hand is so old school, unc.”

quote

Don’t get me wrong, I’m not against writing code with AI. I welcome it as a tool to get things done faster (debatable in some scenarios).

I’m also not going to deny that every now and then, I miss the rush I felt every day in my head and my fingertips as they hit the keys - that rythm and those sweet clicks - in a flow state that made me feel like a wizard and an imposter at the same time.

I guess I will probably save some money keyboard shopping (also debatable).

I’ve read a lot, though.

Diffs, docs, chat transcripts. My own writing from a few weeks ago that I’m now double-checking because an agent helped me draft it and I never gave it the second pass it deserved.

I don’t have a clean name for the specific tired that comes from a day like that.

It’s not the tired from writing something hard. It’s closer to what people are starting to call AI agent fatigue — or atleast that’s what the internet bubble I live in has led me to believe — and it turns out that’s a more literal, more measurable thing than I expected when I sat down to write this.

Yes, I took out time to read more stuff so I can write about how exhausted I am reading stuff. I mean.. I chose this madness when I thought engineering is cool.

This Isn’t the Multitasking We Already Know

The easy explanation is that I’m just bad at multitasking like everyone else (basic human neurology and what not), or that context-switching is expensive and always has been..

Both of those are true but that doesn’t seem like it is the whole story.

The context-switching cost predates any of this. From what I found — read: Google-searched and clicked on reference links Gemini gave me and then went down the rabbit hole with whatever I had left in me — a 2008 study by Gloria Mark, Daniela Gudith and Ulrich Klocke(opens in new tab) on interrupted work found something that surprised me when I ran into it.

People actually finished interrupted tasks in slightly less time than people working without interruption (about 20.3 to 20.6 minutes versus 22.8 minutes). Sounds good, huh?

But they did it under measurably more stress, frustration, and time pressure. And the study showed that it was independent of the personality (like politness and openness). They also saw that their emails got shorter, which honestly I don’t mind.

The more important point is that it suggests people compensated for the time they expect to lose by working faster and writing less at the price of more stress. Even a recovery delay would probably mean people try to make up for it in a stressful pace.

quote

A separate study the google search result led me to found it takes roughly 11 to 16 minutes to fully resolve an interruption before you’re back on the original task. (Iqbal & Horvitz, CHI 2007.)

None of that is new, and none of it is about AI. It’s just the baseline cost of a human brain switching lanes. I think what’s happening across five parallel AI sessions is a different shape of the same problem.

Ordinary multitasking is switching between things you are doing.

This is switching between things something else already did, where the only work left for me is deciding whether to trust it.

That’s a supervisory job, not a “doing” job, and it doesn’t parallelize the way generation does. Five agents can write code at the same time. I can only decide, one at a time, whether each of them got it right.

It also means there’s no flow state to protect because there was never one to begin with. Flow needs sustained attention on a problem you’re building. Watching a queue of finished-looking work and deciding yes or no, five times over, is closer to designed interruption than to deep work.

There’s nothing to fall out of, which somehow makes it worse..

The Burnout Risk has probably Increased

I find myself circling back to one more thought today.

Producing the code got faster. Deciding whether to trust what got produced did not get faster, and I don’t think it can, because trust gets evaluated one thread at a time, in one head.

And I don’t think that’s just how I feel. Sonar’s 2026 developer survey(opens in new tab), covering over 1,100 professional developers, found that 96% of them don’t fully trust AI-generated code. I wasn’t part of that survey but you know I relate with the 96% and already feel bad for the other 4%.

Only 48% said they always verify it before committing anyway. 💀

That tells me distrust isn’t converting into a habit of actually checking, which is its own quiet problem sitting underneath the bigger one.

bonus

Qodo’s 2026 State of AI Code Quality report(opens in new tab), based on 500 developers and 300 engineering leaders, put a name on the specific pattern: roughly a third of developers describe a “hidden trust tax,” where reviewing AI-written code takes about the same amount of time it always did but demands noticeably more cognitive effort to catch the subtle stuff.

Review and validation came out as the single most-cited bottleneck in software delivery, on both the developer side and the leadership side.

That’s exactly what today looked like.

The PR sitting in my queue isn’t hard because the code is bad.

It’s hard because I have to actually think about it, the way I would if a colleague wrote it, except there are four other things waiting for the same kind of thinking. And 38% of the folks from that Sonar survey admitted that it is harder to review code written by AI compared to their human colleagues.

I also want to be straight about this: both of those surveys come from vendors who sell tools into this exact problem, so I’m treating the numbers as directional, not gospel.

But between the two of them, distrust that isn’t turning into verification on one side, and review time staying flat while review effort climbs on the other, they’re both pointing at the same general spot.

Trust and review, not code generation, is where the friction actually sits.

quote

I also wrote a short rant about this gap — the difference between code that looks finished and code a reviewer has actually verified — in a bit about broken access control slipping past AI-assisted reviews. That habit works fine when you’re applying it once. It gets a lot harder to hold onto when you’re applying it five times before lunch, on five different mental models, each one context-switched cold.

The Industry’s Answer Is More Agents, Not Less Watching

Well.. 🤦🏻

While I was looking into all this, I ran into GitHub’s own recent push in the opposite direction. The Copilot app now lets you run several agent sessions at once(opens in new tab), each on its own isolated git worktree, so nothing steps on anything else.

Copilot CLI’s /fleet command, shipped earlier this year(opens in new tab), does something similar — splitting one task into independent pieces and running them in parallel. I had already tried this on my personal Claude subscription by instructing my agents to create worktrees and work off them.. for funsies. I could easily imagine the mental cost of running it every day over and over again.

GitHub’s own framing for the desktop app is that it frees you up to spend your time “reviewing and making decisions rather than watching your agents at work.”

I don’t think that’s dishonest marketing, exactly. But read it again next to everything above.

Reviewing and deciding is the bottleneck this whole article is about. Making it easier to start a fourth or fifth session doesn’t touch that bottleneck.

It just makes it easier to reach. 🤷🏻‍♂️

What Actually Helped, a Little

I did solve the bug today and handed it off to a real human colleague (yes, thankfully we still have those and I’m very grateful for it), skimmed the design doc properly on my second real pass, and got through the PR review by forcing myself to stop toggling and take the sessions one at a time instead of glancing between all of them every few minutes.

That’s not a framework.

It’s just the one thing that made today feel less like triage: treating review as its own dedicated block of attention instead of something I did in the gaps between checking on agents.

It didn’t make the work smaller. It made the part I was actually bad at today - batching my own judgment instead of fragmenting it - slightly less bad. I still don’t have a tidy name for the tired I felt by the end of the day. But I know now it’s not just me, and it’s not just ordinary multitasking either.

It’s a newer, narrower kind of work: five things producing, one person deciding, and no one ever sat us down to explain how that part was supposed to go..