Whole-system performance for agentic hardware development

Your agents can build the system. Kairolith tells them how it performs.

Kairolith is a whole-system performance tool for hardware teams that build with coding agents. It runs your real software on a simulated version of your complete system, including the chips still in design, and reports end-to-end performance and the component responsible on every change.

Coding agents now write drivers, firmware and chip designs, and every piece passes its tests. What nobody sees yet is how the finished system will perform. Kairolith measures that on every change, long before there is silicon, and names the component that costs the time. Problems show up while they are still cheap to fix.

click to play with sound
Read the video as text · 46 s

Coding agents write the RTL, the firmware and the driver, and every piece passes its own tests, but nobody can tell how the assembled system performs. The video shows Kairolith running the hardware and the unmodified software together on one clock, breaking a 178 µs round trip down by component, and an agent finding a 12 µs wait on the network card, fixing it and getting back exactly what changed, long before silicon.

  1. Agents build it

    Your agents can write the RTL. The firmware. The driver.

  2. Every piece passes

    Every piece passes. Then the pieces meet, and nobody can tell you how the system performs.

  3. The request path

    The number that matters is the path a request takes, across all components in the system. No component test measures it.

  4. One clock

    Kairolith runs the whole thing. Your hardware, your unmodified software, on one clock. And shows you where the time goes.

  5. The loop, live

    While the run is still going, your agent sees the first round trips, finds twelve microseconds waiting on the NIC, edits, resubmits, and gets back exactly what changed.

  6. Long before silicon

    Hardware development, as a software loop. Long before silicon. One link. Paste it into your agent.

The problem

Agents speed up every part. The system is where the risk is.

Specialized hardware only pays off if the whole system is fast: your chip, the driver and the software around it, working together. That is the number your customers buy and the number your roadmap depends on.

No single part can tell you that number. Tests and benchmarks check one component at a time. The system result arrives once the hardware is built, and by then a change to the design costs the most.

Coding agents make this sharper. They produce more changes, faster, to more parts of the system. Each change passes its tests. What it did to the performance of the whole is unknown, and agents cannot check it on their own.

each part
Tested

Unit tests, testbenches and benchmarks pass for the driver, the firmware and the chip design.

each agent change
Correct

It builds, it passes review, it does what the ticket asked for.

the whole system
Unknown

How fast it is, and which part is holding it back, stays open until there is hardware to measure.

What you get

A performance answer on every change, while the design can still change.

01 · earlier

Know if you will hit the target, before tape-out

Kairolith runs your real software on a simulated version of your system, including the parts you are still designing. You see end-to-end performance while the driver, the firmware and the chip design can all still change.

02 · why

Every result names the part responsible

Not just "178 µs", but "12 µs of it is the server waiting on the network card". The fix goes to the right team the first time, instead of starting a hunt across groups.

03 · agents do the loop

Your engineers set targets. Agents do the iterations.

Run, measure, find the cause, try again: that loop is what agents are good at once they can see the result. Your engineers decide the targets and review what actually moved the number.

04 · guarded

A slowdown shows up on the commit that caused it

Performance targets run next to your existing tests. When a change makes the system slower, you find out on that change, with the component named, not weeks later.

An example

"Make the round trip faster." What happened next.

  1. The request. A message between two servers, through your network cards and a switch, takes 178 µs. The target is 170. An engineer asks the agent to fix it.
  2. The measurement. The agent starts the whole system in Kairolith. Three minutes in, with the run still going, it already has the answer.
  3. The cause. 12 µs of every round trip is the server waiting on its network card, which holds messages back on a timer.
  4. The fix. The agent changes one setting in the driver and runs again. Kairolith confirms that only the driver changed and the wait is gone.
Round trip178 → 166 µs
Hardware builtNone
Engineer timeOne request
one round trip · time per componentwhat the agent sees
client serveryour softwareunmodified card Ayour design switchstandard part card Byour design serveryour softwareunmodified cause: card B holds the message back client · 41 server · 34 reply · 78 client · 21 12 µs waiting 0 100 178 µs
Example system. The first-run timings are real; the second run is illustrative.

Two fair questions

Why not use what we already have?

"We already have simulators and tests. Our agents can run those."

They can, and each of those runs is right about its own part. The chip simulation passes, the software model passes, the network model passes. None of them contains the 12 µs, because the delay only happens when your software, your chip design and the rest of the system run together, in step.

Kairolith builds that combined system from the models and simulators you already have, so nothing gets thrown away. It then hands the agent a short answer in your own terms, instead of piles of logs from separate tools.

"Our agents are good. Can't they just estimate the effect?"

They can explain any change convincingly. Whether it actually helps depends on how it interacts with everything else in the system, and that is hard to predict from the code. In the example system an agent proposed three changes, each with a good reason:

bigger message queueslowest requests 9 µs slower
shorter interrupt timerround trip 12 µs faster
earlier data prefetchround trip 4 µs faster

All three sounded right. One made things worse. You only know which by measuring, the same way code is not merged because the agent says it works.

Across the team

One tool, three kinds of answer.

Engineers

Ask in plain language why something got slower. The agent runs the system, finds the cause and proposes a fix, with the measurement attached for review.

Engineering leads

Performance targets are checked on every commit, like tests. A regression is caught on the change that caused it, with the responsible component named.

Architects and leadership

A running record of which changes moved the number and why, across everything the agents tried. Realistic system numbers long before hardware, to plan on and to show customers.

Get started

Start with one system and one target.

  1. Pick a system you are buildingWe set up Kairolith with your team, using your software and the models you already have for your own hardware.
  2. Set the targetFor example, a round trip under 170 µs. From then on every change is measured against it.
  3. Let the agents workYour agents iterate against the target. Your engineers review what moved the number, and why.

For your engineers: Kairolith connects to Claude Code, Codex or your own agent with one line. The agent reads the instructions and sets itself up.

Read https://[host]/agent.md and follow it.

Frequently asked questions

What teams ask before a pilot.

Can a faster chip make the system slower?

Yes. In one system we measured, an accelerator that was 25% faster on its own lost end to end once the link to the host took more than about 0.7 µs, because its driver waited on that link 64 times per request.

Does my design have to be in RTL?

No. Bring the models you already have: RTL through Verilator, SystemC, or a behavioral model in C++, and switch to RTL later without changing the rest. Hosts, network and standard devices come with Kairolith.

Do we have to change our software?

No. The kernel, drivers, firmware and applications run unmodified. Linux works out of the box, and your driver on top of it needs nothing extra.

How accurate are the numbers?

Every component runs on one synchronized clock, and runs are reproducible, so every difference between two runs is real, not noise. Each result names the model behind each component, so you know whether you measured your design or a stand-in.

Isn't full-system simulation slow?

A run takes longer than real hardware, but you don't wait for the end. Results arrive while it is still going, and you or your agent stop when you have seen enough.

Does our design leave our network?

It doesn't have to. Kairolith runs on your own infrastructure or in the cloud.