From the field

What I'm seeing building AI agents across 4 ventures in 2026

Julien de Waal Sep 19, 2026 7 min read

Not theory. Notes from shipping agentic systems in music, content, and SaaS this year — what works, what breaks, and what keeps pulling the timeline forward.

I've been running agentic systems across four ventures simultaneously for most of 2026. Music video production. Social content distribution. An AI music label. And this site, which exists because I needed a reliable source for AGI forecast data that wasn't wrapped in vendor spin.

This isn't a think-piece about AI agents. It's what I've actually observed — the patterns that hold across domains and the friction that's still real.

What's working

Agents that stay in their lane

The highest-ROI use case, consistently, is agents with a narrow job and a clear exit condition. An agent that monitors competitor pricing and surfaces anomalies. An agent that takes a piece of content and adapts it for three platforms. An agent that scans incoming signals and routes them to the right queue.

These aren't glamorous. They're unglamorous and they just work. The agents that fail are the ones trying to do too much in a single pass — make a judgment call, take an action, evaluate the result, loop. That chain breaks. Narrow scope, clear handoff.

Field note — Sprinkal
Social content distribution

The distribution agent does one thing: take a finished piece and schedule it across platforms with format-appropriate adaptation. It doesn't decide what's worth distributing. That stays human. When we tried to add a "should we post this?" judgment layer, the agent became unreliable. Split the jobs. Ship faster.

The context window is the real constraint

Everyone talks about capability. The real daily constraint is context. A well-prompted agent with a rich context window outperforms a poorly-prompted agent with a capable model. Every time. Getting the context right — what does this agent need to know to do its job well — is 80% of the work.

This is also why RAG pipelines matter more than most people outside the field appreciate. The agent that can pull the right piece of context at the right moment is dramatically more useful than one relying purely on what's in the prompt.

Agentic systems compound

The second and third month of running an agentic system is better than the first, because the system has more examples to learn from — not in a training sense, but in a prompt-tuning, edge-case-handling sense. You find the failure modes. You patch the prompts. You add guardrails around the things that break. The system gets tighter.

This is the compounding argument for building early. The agent you've been running for six months is not the same agent you launched. It's been shaped by every edge case you've hit.

The meta-observation Building agentic systems is more like managing a junior team member than like writing software. You set expectations clearly, you give good examples, you review the output, you course-correct. The skills that make a good manager translate surprisingly well to agentic system design.

What's still breaking

Multi-step reasoning that requires external state

Anything that requires the agent to hold a mental model of something external — a relationship, a negotiation, a project's current status — and update it across multiple turns still fails more than it should. The agent loses the thread. It contradicts itself. It treats a follow-up as a first contact.

The fix is usually architectural: give the agent persistent memory that it can read and write explicitly, rather than relying on context alone. This works. It's just more engineering than most demos show.

Field note — Waalhalla Records
AI music label

Artist communication is the hardest thing to hand to an agent. The relationship context — what we discussed three weeks ago, what the artist is sensitive about, where we are in the release cycle — needs to be explicit and retrievable, not implicit. We built a lightweight CRM layer for this. The agent reads it before every interaction. Failure rate dropped significantly.

Creative judgment at the frontier

For tasks where "good" is obvious in retrospect but hard to specify in advance, agents still underperform relative to a skilled human with taste. Music video creative direction. Campaign concept selection. The things where you know it when you see it.

This isn't a reason to not use agents in creative work. It's a reason to keep the taste layer human. The agent can generate 20 options. The human picks one. That workflow is dramatically faster than the human generating from scratch, and the agent's 20 are better than they were a year ago.

Tool reliability at scale

When you're running multiple agents across multiple ventures, tool reliability becomes a systemic issue. An API that's down. A webhook that mis-fires. A third-party rate limit hit at 2am. These are solvable with monitoring and retry logic, but they take real engineering time. The "just connect some agents" abstraction breaks down when you need SLA-level reliability.

The reliability pattern that works Every agent action that touches external state needs: a timeout, a retry with backoff, a failure notification, and a human-readable log entry. This isn't glamorous. It's table stakes for anything you want to trust in production. Build it in from the start, not as an afterthought.

What this tells me about the AGI timeline

Running agents daily gives you a different view of the capability curve than reading benchmarks. The benchmark data says we're on track for AGI around 2031. The field data mostly confirms it, with nuance.

The things that feel close to solved: language tasks, retrieval, narrow execution, format conversion, summarization, first-draft generation. These have crossed a threshold in the last eighteen months where you can build production systems around them without constant babysitting.

The things that still feel genuinely hard: persistent reasoning across time, reliable multi-agent coordination, creative judgment, and anything requiring real-world physical grounding. These are the remaining gaps between current systems and something that deserves the label "general."

The progress rate on the first category has been faster than most people expected. The second category is harder to measure because the goalposts are less clear. But watching it from here, month over month — the trend is consistent. The benchmark curve isn't lying.

The practical implication

If you're building software right now and not thinking seriously about where agents fit in your product — what they can handle, where the human stays in the loop, how the architecture changes — you're accumulating technical debt on a timeline that is compressing faster than most organizational planning cycles.

That's not hyperbole. It's what shipping looks like from inside the curve.

J
Julien de Waal Building AI-native ventures and tracking the AGI timeline closely. Founder of One Person Unicorn — the thesis that the right AI stack changes what's possible for a single operator. Track the live AGI forecast at howcloseisagi.com.

Related