About Surface Area
We built this for F500s before we were acquired by one
Our mission
Testing how an agent will behave in production is dangerous. A wrong move there is a real refund, a real escalation, a real customer. But you cannot know how it behaves without running it, thousands of times, against the systems it will actually touch. So we build somewhere else to run it: a reliable, simulated world.
A world is your production stack rebuilt as something an agent can act on: your data mirrored under the governance it already lives under, your tools and APIs with the failure modes they actually have, your policies as limits it can genuinely violate. Where the data cannot leave your environment we backfill it with generated records that hold the same shape and edge cases. The build runs as a fixed sequence with tests at every stage, because a world that drifts from production hands you numbers that look fine and describe a system you do not run.
What clears the world is what ships. None of it has to stay internal either: the environments and scoring you build to prove your own systems can be published as benchmarks under your name.
Our team
The founders built eval infrastructure and agents across multiple Fortune 500 enterprises and frontier labs. One of those enterprises acquired them, and they went on to lead Applied AI across the business. The broader team comes from Stripe, Notion, Meta, and frontier AI labs.