WEBVTT

00:00:00.000 --> 00:00:03.200
Most of us are stuck in a basic AI loop. Yeah,

00:00:03.299 --> 00:00:06.139
that endless hamster wheel. Exactly. You write

00:00:06.139 --> 00:00:08.320
a prompt, you get a bad result, and you just

00:00:08.320 --> 00:00:11.160
keep tweaking that same prompt. That is so frustrating.

00:00:12.099 --> 00:00:14.820
And notoriously slow. But there is a much more

00:00:14.820 --> 00:00:18.739
powerful way. Beat. Welcome to our deep dive

00:00:18.739 --> 00:00:22.359
for today. We are unpacking a workflow from Maestro's

00:00:22.359 --> 00:00:25.519
DayIA. They use Claude code and a tool called

00:00:25.519 --> 00:00:28.510
Ultracode. The mission here is running multiple

00:00:28.510 --> 00:00:32.009
AI agents all at once. Yeah, this workflow completely

00:00:32.009 --> 00:00:35.170
flips the standard script. We are moving way

00:00:35.170 --> 00:00:37.570
beyond basic prompt engineering today. Right.

00:00:37.630 --> 00:00:39.490
We're going to cover the foundational harness

00:00:39.490 --> 00:00:41.950
structure first. We'll look at these intense

00:00:41.950 --> 00:00:45.570
verification loops, and then we'll see true parallel

00:00:45.570 --> 00:00:48.009
processing in action. To break out of that simple

00:00:48.009 --> 00:00:50.210
prompt cycle, things must change. We have to

00:00:50.210 --> 00:00:52.310
change the environment the AI operates in. You

00:00:52.310 --> 00:00:54.310
can't just keep throwing tasks into a blank chat

00:00:54.310 --> 00:00:56.909
box. Right. Which brings us to a concept called

00:00:56.909 --> 00:00:59.170
the harness. Let's clearly define that technical

00:00:59.170 --> 00:01:02.170
term right now. A harness is a setup of rules

00:01:02.170 --> 00:01:05.069
and tools guiding an AI's specific job. That

00:01:05.069 --> 00:01:07.989
is a highly accurate way to frame it. You are

00:01:07.989 --> 00:01:10.650
giving the agent a very restricted workspace.

00:01:11.569 --> 00:01:14.489
You remove the chaos of an open -ended conversation

00:01:14.489 --> 00:01:16.909
completely. It is like hiring a brand new employee.

00:01:16.969 --> 00:01:19.010
You don't just yell a task at them in the hallway.

00:01:19.090 --> 00:01:21.090
Right, you'd never do that. You give them a full

00:01:21.090 --> 00:01:23.349
onboarding packet. You give them a dedicated

00:01:23.349 --> 00:01:26.480
desk and resources. Aaron Levy actually highlights

00:01:26.480 --> 00:01:29.019
this exact concept. Oh, yeah. He says an agent

00:01:29.019 --> 00:01:31.920
harness is crucial for complex work. Right. And

00:01:31.920 --> 00:01:34.280
without that structure, agents just get completely

00:01:34.280 --> 00:01:37.400
lost. They desperately need those rigid boundaries.

00:01:37.840 --> 00:01:41.079
Otherwise, the AI tries to be the researcher,

00:01:41.379 --> 00:01:43.939
writer, and editor simultaneously. I still wrestle

00:01:43.939 --> 00:01:46.560
with prompt drift myself. Yeah. Two -sec silence.

00:01:47.000 --> 00:01:49.900
I often start with a very clear goal. The AI

00:01:49.900 --> 00:01:52.819
gives me a decent draft, but I want one small

00:01:52.819 --> 00:01:55.510
change. I know where this is going. I ask it

00:01:55.510 --> 00:01:58.489
to adjust one specific paragraph. It fixes it,

00:01:58.489 --> 00:02:00.329
but somehow ruins all the formatting. Exactly.

00:02:00.370 --> 00:02:02.829
After three prompts, it has forgotten my very

00:02:02.829 --> 00:02:05.870
first instruction entirely. Breaking instructions

00:02:05.870 --> 00:02:09.030
into smaller, harnessed skills fixes this. And

00:02:09.030 --> 00:02:11.310
that focused environment leads right into the

00:02:11.310 --> 00:02:13.830
verification loop. The workflow does not just

00:02:13.830 --> 00:02:16.370
blindly accept the first answer. Right. Standard

00:02:16.370 --> 00:02:18.849
AI outputs text and just assumes it's finished.

00:02:19.270 --> 00:02:21.949
But this system adds a mandatory quality control

00:02:21.949 --> 00:02:25.699
checkpoint. Addy Osmani calls this loop engineering.

00:02:26.060 --> 00:02:28.740
OK, loop engineering. Yeah. Outputs are scored

00:02:28.740 --> 00:02:31.340
against very strict criteria. If the score is

00:02:31.340 --> 00:02:34.319
too low, the task simply runs again. So it actively

00:02:34.319 --> 00:02:37.599
refuses to move forward with bad data. Precisely.

00:02:37.939 --> 00:02:40.219
It keeps looping back until it passes the test.

00:02:40.479 --> 00:02:42.340
But wait, how does the workflow prevent an endless

00:02:42.340 --> 00:02:45.479
loop if the output keeps failing? It relies on

00:02:45.479 --> 00:02:48.039
highly specific checking criteria to eventually

00:02:48.039 --> 00:02:52.080
force a pass. The rubric asks binary yes or no

00:02:52.080 --> 00:02:55.180
questions to break the cycle. So strict grading

00:02:55.180 --> 00:02:57.680
rubrics force the agent to eventually hit the

00:02:57.680 --> 00:03:00.080
required target. Exactly. It creates a mathematically

00:03:00.080 --> 00:03:02.800
guaranteed baseline of quality. But this verification

00:03:02.800 --> 00:03:05.159
loop is only effective under certain conditions.

00:03:05.800 --> 00:03:08.080
The automated grader cannot be biased by the

00:03:08.080 --> 00:03:10.840
worker's messy process. This is exactly why fresh

00:03:10.840 --> 00:03:13.560
context is so incredibly important. The verifier

00:03:13.560 --> 00:03:15.960
agent is kept entirely separate from the worker

00:03:15.960 --> 00:03:19.580
agent. Right. It doesn't carry over any context

00:03:19.580 --> 00:03:22.300
from the initial work. It gets a completely wiped

00:03:22.300 --> 00:03:25.060
blank slate memory to work with. Two sec silence.

00:03:25.780 --> 00:03:28.360
That architectural separation makes a massive

00:03:28.360 --> 00:03:31.180
amount of sense. It is like having an independent

00:03:31.180 --> 00:03:34.280
auditor check a company's financials. Yeah. Perfect

00:03:34.280 --> 00:03:36.259
analogy. You don't ask the accountant who made

00:03:36.259 --> 00:03:38.819
the errors to double check their own math. You'd

00:03:38.819 --> 00:03:41.060
never trust that. Right. You bring in an external

00:03:41.060 --> 00:03:44.680
reviewer with a clean slate. Think about a quality

00:03:44.680 --> 00:03:47.620
control laser on an assembly line. The laser

00:03:47.620 --> 00:03:51.080
only scans the final car door for dents. It doesn't

00:03:51.080 --> 00:03:53.139
ask the welding robot how it feels about its

00:03:53.139 --> 00:03:55.659
job. Yeah, that makes sense. It objectively measures

00:03:55.659 --> 00:03:57.439
the physical product against the blueprints.

00:03:57.879 --> 00:04:00.539
This matches finder and verifier setups across

00:04:00.539 --> 00:04:03.120
all of AI. The verifier just objectively looks

00:04:03.120 --> 00:04:06.060
at the final output. The context window isn't

00:04:06.060 --> 00:04:08.280
cluttered with previous failed attempts. But

00:04:08.280 --> 00:04:10.159
this still raises an interesting architectural

00:04:10.159 --> 00:04:13.199
question. Why can't the original worker agent

00:04:13.199 --> 00:04:16.139
just be told to check its own work? Because the

00:04:16.139 --> 00:04:18.920
worker carries the heavy baggage of its past

00:04:18.920 --> 00:04:22.220
decisions, it naturally assumes its original

00:04:22.220 --> 00:04:24.839
internal logic is entirely correct. Because the

00:04:24.839 --> 00:04:26.920
worker's hidden assumptions would blind it to

00:04:26.920 --> 00:04:29.680
its own mistakes. Precisely. you desperately

00:04:29.680 --> 00:04:32.699
need a fresh pair of digital eyes. So we have

00:04:32.699 --> 00:04:34.920
this high quality verification system in place.

00:04:35.100 --> 00:04:37.399
Right. We have a worker and a fresh verifier

00:04:37.399 --> 00:04:40.120
checking it, but doing this sequentially is just

00:04:40.120 --> 00:04:42.199
far too slow. Oh, incredibly slow. Waiting for

00:04:42.199 --> 00:04:45.319
each step creates a massive operational bottleneck.

00:04:45.699 --> 00:04:48.800
We need to speed this complex process up significantly.

00:04:49.699 --> 00:04:52.699
This brings us to a concept called graph engineering.

00:04:52.879 --> 00:04:55.060
Yeah, this is the really exciting part. Let's

00:04:55.060 --> 00:04:58.350
define graph engineering clearly. It is splitting

00:04:58.350 --> 00:05:00.670
independent tasks into separate paths that run

00:05:00.670 --> 00:05:02.970
at the same time. This is where the Maestro's

00:05:02.970 --> 00:05:05.889
DIA workflow gets seriously impressive. Each

00:05:05.889 --> 00:05:08.629
independent task gets its very own branch. They

00:05:08.629 --> 00:05:11.310
don't have to wait in a slow line anymore. Let's

00:05:11.310 --> 00:05:14.129
look closely at the specific demo they ran. They

00:05:14.129 --> 00:05:17.009
used Claude code and Ultra code to build this

00:05:17.009 --> 00:05:21.290
structure. Just typing slash Ultra code and slash

00:05:21.290 --> 00:05:23.350
workflows into the command line. The primary

00:05:23.350 --> 00:05:26.139
mission was to research three different YouTube

00:05:26.139 --> 00:05:30.439
channels, OpenAI, Anthropic, and Google DeepMind.

00:05:30.620 --> 00:05:33.000
And they specifically wanted videos published

00:05:33.000 --> 00:05:35.300
in the last seven days. Think about doing this

00:05:35.300 --> 00:05:38.040
normally. You would wait for the OpenAI research

00:05:38.040 --> 00:05:40.899
to finish completely, then Anthropic would slowly

00:05:40.899 --> 00:05:44.459
take its turn. But here, three separate AI agents

00:05:44.459 --> 00:05:47.259
spin up and run simultaneously. Instead of waiting

00:05:47.259 --> 00:05:49.300
patiently, the system branches out immediately.

00:05:49.600 --> 00:05:51.740
It is like cooking a massive dinner for a large

00:05:51.740 --> 00:05:53.779
party. You don't wait for the chicken to finish

00:05:53.779 --> 00:05:56.740
before chopping carrots. Exactly. Akshay Pachar

00:05:56.740 --> 00:05:59.439
has actually demonstrated very similar subagent

00:05:59.439 --> 00:06:02.709
splits recently. Paul Huron also shows how context

00:06:02.709 --> 00:06:05.990
is isolated to its specific task. That isolation

00:06:05.990 --> 00:06:09.550
is the absolute magic of true parallel execution.

00:06:09.870 --> 00:06:11.990
And the strict verification loop still happens

00:06:11.990 --> 00:06:14.990
inside each branch. OK, so the verifier agents

00:06:14.990 --> 00:06:17.829
review outputs without talking to the other branches.

00:06:17.850 --> 00:06:20.029
Yeah, entirely separately. Once all branches

00:06:20.029 --> 00:06:22.589
are finished and verified, things finally merge.

00:06:23.089 --> 00:06:25.610
Claude Code brings everything into one final

00:06:25.610 --> 00:06:30.009
comparative PDF report, beat. But managing all

00:06:29.839 --> 00:06:32.639
these streams seems precarious. Does running

00:06:32.639 --> 00:06:34.959
these branches at the same time mix up their

00:06:34.959 --> 00:06:38.060
data? Not at all. The underlying system walls

00:06:38.060 --> 00:06:40.920
off the context for each branch entirely. They

00:06:40.920 --> 00:06:42.800
don't share any memory until the final merge.

00:06:43.019 --> 00:06:45.279
Got it. Each branch keeps its context strictly

00:06:45.279 --> 00:06:48.110
isolated until the final combination. That strict

00:06:48.110 --> 00:06:50.970
isolation makes the parallel processing so incredibly

00:06:50.970 --> 00:06:53.610
reliable. You don't get agents hallucinating

00:06:53.610 --> 00:06:55.730
details across different companies. We understand

00:06:55.730 --> 00:06:57.649
the theory behind this complex structure now.

00:06:57.689 --> 00:07:00.189
Yeah. But what exactly happened when they hit

00:07:00.189 --> 00:07:03.250
enter on this workflow? What did the real world

00:07:03.250 --> 00:07:05.829
execution actually look like? The initial demo

00:07:05.829 --> 00:07:08.189
results are honestly pretty mind blowing. It

00:07:08.189 --> 00:07:10.569
was incredibly fast. How fast? Each research

00:07:10.569 --> 00:07:12.649
session took only about 40 seconds to finish

00:07:12.649 --> 00:07:16.069
and used roughly 54 ,000 tokens to process the

00:07:16.069 --> 00:07:18.970
information. That is a truly massive amount of

00:07:18.970 --> 00:07:22.769
raw data. Processing 54 ,000 tokens is roughly

00:07:22.769 --> 00:07:25.889
equivalent to reading a short novel. And the

00:07:25.889 --> 00:07:28.209
system processed that in under a minute. Whoa,

00:07:28.389 --> 00:07:31.370
imagine scaling to a billion queries. The sheer

00:07:31.370 --> 00:07:34.370
volume of automated processing is completely

00:07:34.370 --> 00:07:37.569
staggering. Two -sec silence. Let's talk about

00:07:37.569 --> 00:07:40.879
the verification mechanics in action. What exactly

00:07:40.879 --> 00:07:43.360
did these independent verifier agents catch?

00:07:43.600 --> 00:07:45.860
They caught some very critical timeline errors

00:07:45.860 --> 00:07:48.579
immediately. None of the five initial videos

00:07:48.579 --> 00:07:50.360
were actually published recently. We know them.

00:07:50.519 --> 00:07:52.939
Nope. They fell way outside the strict seven

00:07:52.939 --> 00:07:55.199
-day window. So the initial worker agents just

00:07:55.199 --> 00:07:58.000
completely hallucinated the publication dates.

00:07:58.500 --> 00:08:00.579
Exactly. Language models are notoriously bad

00:08:00.579 --> 00:08:03.339
at strictly tracking chronological time, but

00:08:03.339 --> 00:08:05.660
the fresh verifiers caught that instantly. They

00:08:05.660 --> 00:08:08.220
also flagged a highly specific Google DeepMind

00:08:08.220 --> 00:08:10.779
video. It could not be confirmed on the official

00:08:10.779 --> 00:08:13.259
channel at all. So the verification loop actively

00:08:13.259 --> 00:08:15.939
stopped bad data, but they also confirmed plenty

00:08:15.939 --> 00:08:18.459
of good data during this. Oh, absolutely. The

00:08:18.459 --> 00:08:21.279
verifications confirmed six videos with zero

00:08:21.279 --> 00:08:23.920
discrepancies whatsoever. It really showcases

00:08:23.920 --> 00:08:26.779
the rigorous nature of this setup. They validate

00:08:26.779 --> 00:08:29.160
the good data while catching the subtle errors.

00:08:29.420 --> 00:08:32.580
all clearly displayed in the final PDF. There

00:08:32.580 --> 00:08:35.059
is almost always a hidden catch with these powerful

00:08:35.059 --> 00:08:38.419
systems. I have to note the very real operational

00:08:38.419 --> 00:08:40.960
trade -off here. Yeah, the costs. Adding separate

00:08:40.960 --> 00:08:43.679
workers and verifiers burns through tokens incredibly

00:08:43.679 --> 00:08:46.220
fast. You are essentially paying for three or

00:08:46.220 --> 00:08:49.200
four API calls instead of one. It is definitively

00:08:49.200 --> 00:08:51.779
a much more expensive way to operate. For an

00:08:51.779 --> 00:08:54.299
enterprise company, that token cost is a rounding

00:08:54.299 --> 00:08:57.629
error. But for an individual developer, it really

00:08:57.629 --> 00:08:59.870
adds up. Which brings up a very practical concern.

00:09:00.450 --> 00:09:03.990
With such high token usage, is this setup actually

00:09:03.990 --> 00:09:06.850
practical for everyday tasks? It is strictly

00:09:06.850 --> 00:09:09.350
for large tasks where accuracy is completely

00:09:09.350 --> 00:09:11.590
non -negotiable. You wouldn't use this heavy

00:09:11.590 --> 00:09:14.870
framework to draft a casual email. More tokens

00:09:14.870 --> 00:09:17.990
mean higher costs, but you get guaranteed quality

00:09:17.990 --> 00:09:20.070
in return. That is the ultimate trade -off you

00:09:20.070 --> 00:09:22.289
have to weigh. You are buying structural peace

00:09:22.289 --> 00:09:24.370
of mind. Let's step back for a moment and summarize

00:09:24.370 --> 00:09:27.129
the core lesson. Let's recap. An AI workflow

00:09:27.129 --> 00:09:30.370
is vastly stronger when each agent has a clear

00:09:30.370 --> 00:09:34.210
role. Tasks run in parallel branches, and outputs

00:09:34.210 --> 00:09:36.590
are checked by an independent verifier with fresh

00:09:36.590 --> 00:09:39.769
context. It totally redefines how we should interact

00:09:39.769 --> 00:09:42.250
with language models. We're moving from being

00:09:42.250 --> 00:09:46.230
prompt writers to actual system architects. But

00:09:46.230 --> 00:09:48.570
understanding this does leave me with a slightly

00:09:48.570 --> 00:09:51.610
provocative thought. If we are building systems

00:09:51.610 --> 00:09:56.409
where AI independently verifies other AI, what

00:09:56.409 --> 00:09:58.409
happens when these agents begin writing their

00:09:58.409 --> 00:10:01.309
own strict grading criteria? Who verifies the

00:10:01.309 --> 00:10:03.690
verifier when the entire system becomes fully

00:10:03.690 --> 00:10:06.429
autonomous? That is a deeply fascinating problem

00:10:06.429 --> 00:10:09.100
for the very near future. Thank you for taking

00:10:09.100 --> 00:10:11.000
this deep dive with us today. We highly encourage

00:10:11.000 --> 00:10:13.379
you to experiment with workflow branching yourself.

00:10:13.860 --> 00:10:16.259
Stop tweaking the exact same basic prompt over

00:10:16.259 --> 00:10:18.460
and over, build a strict harness, add a clean

00:10:18.460 --> 00:10:20.919
verifier, and run it all in parallel. Out to

00:10:20.919 --> 00:10:21.360
your own music.
