Infinitely customizable
evals tooling for robotics
SignalFlag is the test infrastructure platform for physical AI.
Our MCP turns your agent into a robot-evals expert, building higher-quality test infrastructure
10x faster at a fraction of the cost.
Already building on SignalFlag
The evals foundation
your agent is missing
Build from scratch
Rushed implementation
‘Good enough’ quality
Ongoing maintenance
Customizable
Use the SignalFlag App
Works out-of-the-box
Thoughtful UX
Optimized infra
Few custom features
Use the SignalFlag MCP
Rapid implementation
Quality UX components
Optimized infra components
Infinite, stable customization
We maintain the plumbing,
so this stops being your problem...
“There’s no owner for it long-term.
That infra is going to rot pretty fast.”
Every modality, one source of truth
All robotics teams use multiple modalities of testing to evaluate performance.
SignalFlag supports them all, in one place.




Bring your own agent,
your own simulator,
your own CI.
Still your build. Just not from zero.
Connect your agent once. It builds against components already running in production, in whatever shape your team wants.
- 01
Connect
claude mcp add --transport http signalflag https://bff.resim.ai/mcp
- 02
Ask
Set up a regression suite that runs my replay logs on every PR and flags any performance drop.
agent
reading available components…
writing signalflag/evals.yaml
writing regression.yaml
done — 2 files, 31 lines
- 03
Run
signalflag test-suites run pipeline-prod
⚡ [12:04:10] executing 3.5k test cases...
testing complete [3.1k pass, 300 warn, 100 block]
Ready to get started?
FAQs
What is SignalFlag?
SignalFlag is a test infrastructure platform for robotics and physical AI.
Teams connect their coding agent to our MCP, and it builds test suites, metrics, and CI on components already running in production, instead of writing that foundation from scratch.
Field logs, hardware runs, replay, and simulation all land in one place.
SignalFlag was formerly ReSim.
Why do I need SignalFlag if my agent can already build test infrastructure?
Your agent will write you a working test harness in an afternoon. The question is what happens after.
Test infrastructure written from scratch is code your team owns, debugs, and maintains, and it usually only makes sense to the person who prompted it.
Pointed at SignalFlag, the same agent builds against components we already maintain, so what it hands back is shorter, more stable, and consistent with what everyone else on your team is building. The fundamentals are also included, so you can start with a world class evals platform right out of the box as opposed to building it all from scratch.
Do we have to use an agent?
No. Everything reachable through the MCP is also available through the CLI and API, so teams that would rather write it themselves can.
The agent path is the fastest way in, not the only one.
How is this different from a simulator?
SignalFlag runs the tests you define across whatever you already use, including Isaac Sim, Gazebo, MuJoCo, or your own, alongside hardware runs, replay, and field logs.
The value is in orchestrating those runs and making their results comparable, not in producing the simulation itself.
We do not build simulators, and we do not replace the one you have. In this way, we enable you to leverage your own simulator, or depending on the test type, use no simulator at all.
Which agents work with SignalFlag?
Any MCP client. Teams commonly use Claude, Cursor, and Codex.
How does pricing work?
You pay for the compute you use, not for seats, so every plan includes unlimited users. There's a free tier to start, then monthly plans that include usage, with more available as you need it. Larger teams get dedicated pricing.
See the pricing page for details.
Can my team self-host?
Yes. SignalFlag supports on-premises and hybrid deployments, and teams can bring their own storage, including S3 and GCS. If you’d like to meet with a SignalFlag expert about self-hosting, simply .
Does SignalFlag support SSO?
Yes. SignalFlag supports single sign-on through your existing identity provider.
Give your agent a running start
on a foundation built by the experts
Every team builds the same foundation. Yours doesn't have to.







