The problem
Every cloud we rent was designed for people.
The machines are shaped for a human developer: one workstation, kept warm, provisioned for the largest thing it may ever have to do, sitting idle most of the week because the person it belongs to is asleep.
An agent fleet has the opposite shape. It wants twenty machines for two days and none for the next nine. It does not need a desktop, an IDE or a login. It needs a box, the right tools on it, and somewhere to put the result — and then it needs that box to stop existing.
And it does not want one model. A fleet running hundreds of different jobs an hour wants a different model for each of them, chosen on what the job actually needs rather than on which one somebody wired in when the feature was written.
A cloud built for agents, not for people. What Metal is
What it would be
Boxes that appear, work, and go away.
Bare metal we rent by the day, with a model selector in front of it.
-
Dev boxes for agents
A machine an agent can be given for the length of a job: the tools it needs, the repositories it is allowed to touch, and no permanence it has not earned. Rented for a two-day run and released.
-
A model selector
The job says what it needs and the platform decides what runs it. One model is better at writing code, another earns its price only on a genuinely hard problem, and most work needs neither.
-
Local models, local inference
Work that does not need a frontier model should not leave the machine it started on — which is cheaper, faster, and the only honest answer for anything a customer would not want sent anywhere.
-
Rented, not owned
Bare metal by the day. Capacity is something a fleet buys the way it buys tokens, which is the point: more machines is a lever, not a project.
The brain
Find the cheapest model that is still right.
Model choice should be measured, not argued about. This is the mechanism the whole product is built around.
- Take a process somebody does by hand. Answering a class of email, triaging a queue, drafting a particular document.
- Backtest it against reality. Not a dozen examples — five hundred to ten thousand real ones, where the right answer is already known because a person produced it.
- Find the cheapest model that still matches the quality. Not the best model. The cheapest one whose output you would have accepted.
- Turn the answer into a skill the agent selects for itself. The finding stops being a decision somebody has to remember and becomes part of how the fleet routes work.
In our own idea factory, a small model matched the large one at a fifteenth of the cost. Measured, on real examples
Where it stands
The first fleet it has to serve is ours.
We are already doing a version of this by hand, inside our own products, and it already pays for itself. That is the argument for building it properly: the first customer is a fleet we run, so it gets used honestly before it gets sold to anybody.
Metal is a specification and a set of measurements. There is no platform to sign into, nothing to rent, and nothing running behind this page. When that changes, this page changes.
The rest of the engine
One engine, five parts.
You are looking at one of them. They are not a stack you buy in order: each part runs our own companies first, and each one stands on its own. Here are the other four.
-
Agent harness
Airmond Live
Proactive agents that live where you do: Slack, iMessage, email, and on the phone.
-
Agent orchestration
Tartare Runs our own fleet
Every agent we run, ours or a customer’s: runners, credentials, guardrails, dashboards.
-
Agent fleet management
Latch Runs our own fleet
A list of features becomes a sequenced roadmap, and one person runs a fleet of dozens of agents against it.
-
Agent testing
Arena Alpha
Synthetic calls, emails and meetings that test an agent before a customer ever meets one.