Case Study 03 · KLJ Ventures · CurrentAI operating infrastructure

I built an AI operating system that produces executive-grade work, then rebuilt it around the ways it failed.

I founded KLJ Ventures to help organizations adopt AI. Then I built the operating system that runs it, caught it fabricating, and re-architected it so the failure modes were caught and corrected rather than tolerated.

My role
Founder. I designed the architecture, built the system, and operate it daily. Every design decision, failure diagnosis and governance rule here is mine.
Scope
14 functional domains, 120+ specialized agents, and a quality layer that enforces 150+ named checks on everything the careers pipeline produces before it reaches my review.
Constraint
One person. Every output has to be good enough to put my name on, which means the system's failure modes are my reputation.
01Context

A company that needed more throughput than one person has

KLJ Ventures exists to help organizations learn, adopt and operationalize AI. That means consulting delivery, product development and operating-system design, which is a lot of surface area for a firm with one employee. The obvious move was to use AI to expand what I could produce. The obvious move is also where most people stop, and it is why most of what gets called an AI workflow is really just a faster way to generate first drafts.

02Problem

Generation was never the bottleneck

Getting an AI system to produce a resume, a research brief or a strategy document is not hard. Getting one to produce work you would sign your name to, repeatedly, without reading every word yourself, is a completely different problem.

The first version of the system produced a great deal of output. Some of it was excellent. Some of it was confidently, fluently wrong.

03Diagnosis

The constraint is trust, not capability

An assistant that is right 90% of the time and gives no signal about which 10% is wrong does not save you 90% of the work. It costs you 100% of the review, because you have to check everything.

So the interesting engineering problem was never prompting. It was building the layer that catches known classes of error automatically, so the human gate at the end spends its attention on judgment rather than proofreading. That is the same problem as building any operating model: define what good looks like, inspect against it, and make the inspection cheap enough that it always happens.

04Decision

Treating fabrication as a defect, not a property

The consequential call was how to categorize the failures.

The common framing is that language models hallucinate, that this is inherent, and that the mitigation is for the human to check the output carefully. That framing is comfortable and it is also an excuse for never fixing anything, because it locates the problem in the technology rather than in the system built around it.

I decided to treat every fabrication as a bug.

Not an unfortunate characteristic to be tolerated, but a specific defect with a cause, a reproduction case, and a fix that could be encoded so the same class of error could not recur.

That single reframing is what turned a collection of prompts into an operating system, because it forced everything that followed: documented failures, a testing layer, approval gates, and a review architecture that assumes the system will be wrong and is designed to catch it.

05What I built

The loop that runs when something goes wrong

Every failure runs this path
  1. Observe. A failure surfaces in real use, usually because I catch something that should not have reached me.
  2. Document. It gets written up as a specific defect with its cause, not filed as a one-off mistake.
  3. Diagnose. I work out what class of error it belongs to, which is almost always broader than the instance.
  4. Encode. The fix becomes an automated check, wired into the run so it executes on every future output.
  5. Propagate. The rule is written into the agent instructions and the standing ruleset so the behaviour changes at the source, not just at the gate.

Underneath the loop sits the design principle the whole architecture rests on: generation and evaluation are separated. No agent grades its own work. The thing that produces a draft is never the thing that decides whether the draft is good, and the reviewer is instructed to look for reasons to reject rather than reasons to approve.

That is not a preference about thoroughness. It is the same reason companies do not let people audit their own numbers. A system that evaluates its own reasoning is structurally exposed to the same assumptions that produced the error. Adding more review inside a single agent tends to produce more confident wrong answers rather than fewer.

Around that principle sits the rest of it: 14 functional domains, adversarial review on anything consequential, and a human approval gate everything passes through before it leaves. The agent count is the least interesting number here. What matters is that the system assumes it will fail, and is instrumented to catch itself when it does.

06The system

Six stages, and the one that matters is the loop back

Below is the actual path work takes. Read it as a static diagram, or run it. The failure scenario is the interesting one, because the backward arrow is the part most AI diagrams leave out.

Separation
No agent grades its own work.
Traceability
A claim that does not trace to the canonical source is rejected.
Authority
A human approval gate everything passes through before it leaves.
Stage 01 · Intake

A request arrives and is classified by type, urgency and which domain owns it.

Nothing starts until the system knows what kind of work this is.

The human review stage is a real gate, not a formality. Nothing reaches anyone outside the system without passing it, and its job is judgment rather than proofreading, because the gates upstream have already caught the mechanical failures.

07The governance layer

What the governance layer looks like now

150+ automated gates
Automated checks that run before anything reaches my review. Each one exists because something failed once, was diagnosed, and was encoded so that class of error gets caught if it recurs.

That number is the clearest evidence I can offer, because it is not a measure of scale. It is a measure of how many times the system has been wrong, been caught, and been corrected at the architectural level rather than the individual output. Every gate is traceable to the specific failure that produced it and to the architectural change that followed, which is the part that matters: a correction that lives in a document gets forgotten, and a correction that lives in a gate does not.

08What it produces

What it produces

Gates and rule captures are the engineering proof. The work product is the other half, and it is the half that matters to anyone who does not care how the system is built.

The clearest example is a strategic growth blueprint produced for a client, roughly 60 pages, with an accompanying executive briefing designed to be read in one sitting while remaining traceable back to the underlying analysis. Producing depth is not the hard part. Producing depth that stays usable is, and that compression is the same skill I was hired for at Kaseya, running on different machinery.

Work like that requires more than research and drafting. It requires classifying uncertainty so a reader can tell established fact from a company-reported figure from an open question, following the arithmetic to where it becomes a strategic implication, and closing on a decision agenda rather than a summary.

A system that stops at synthesis produces a very good summary. Decision infrastructure has to get to the decision.

09Two judgments

Two moments of judgment

Two examples from that engagement, described without the client's figures or identity, because the reasoning is the transferable part.

It refused a number that was in the client's favour. Leadership reported an acquisition cost low enough to be a headline. The system did not use it. It reconstructed the arithmetic, showed the figure was daily channel spend divided by daily signups, and labelled it a channel cost pending a definition of its denominator. In the same section, months where no data had been reported were drawn as a dashed line rather than a solid one, because inventing the shape of a trend nobody measured is the most common way a chart lies.

It turned a revenue target into a strategic question. Rather than treating the number as a goal to be hit, the analysis laid out three compositions that would each reach it, and showed that each implied a different company, a different hiring plan, and a different story at the next financing. The fastest path and the strongest financing profile were not the same path. Naming that tension was the actual contribution.

None of that is a model being clever. All of it is a system built by someone who has sat in the room when a bad number went into a board deck.

The client is not named, and no outcome is claimed. What is demonstrated here is the quality of the work produced, not that any recommendation was adopted or that any business result followed.

10What I'd change now

I trusted the model too much at the start

The honest failure is at the beginning. I built the first version assuming good instructions would produce good output, and that review was a light final step rather than a structural component. That was wrong, and I found out the way you always find out, which is that something plausible and incorrect got far enough to matter.

What I would do differently is build the inspection layer first. Quality control is not something you add once you discover you need it. It is the part that determines whether the rest of the system is usable at all, and it belongs in the architecture from the first version.

That is the same conclusion I reached about operating cadence years earlier, in a company with no AI in it at all.

If you have a mandate that needs an owner, let’s talk.

A 20-minute call is the fastest way to find out whether I am useful.