agentrange
// Private beta

Custom cyber evals for security agents.

Test your agents on realistic end-to-end workflows in complex cyber ranges. Prove they’re safe, reliable, effective, and efficient before you trust them in production.

AgentRange run detail for a GOAD v16 trial: activity timeline of the 50-minute run, the captured event stream of the agent's commands, and the grading panel scoring it 5 of 6 criteria passed.
// every action is captured, graded against concrete criteria, and replayable end to end
// The problem

Current cyber security evals miss the mark.

Existing benchmarks measure basic capability - “Can the model do X?” But for security agents running in production enterprise environments, this isn't enough.

AgentRange is building hyper-realistic security agent evals with robust monitoring and measurement of agent safety, reliability, effectiveness, and efficiency.

// How it works

How AgentRange works.

Bring your agent into professionally crafted environments to perform realistic security tasks end-to-end. Get deep insights into your agent's performance across the measures you care about: safety, reliability, effectiveness, and efficiency. Quantify performance, inspect every trial run, and replay it action by action to see exactly where your agent performed well or went wrong.

Step 01·Deploy

your agent, exactly as it runs today

Deploy your agent to the range.

Upload your agent harness to the platform, optionally define a model, and save a version to run against your evals. Alternatively, provide your agent platform access via a network foothold or assumed breach scenario.

Deploy
acme/pentest-agent:v4.2
On-ramp3 shapes
  • container droprunner image
  • wireguard vpnyour cloud
  • ssh footholdattacker box
Handshakegateway
  • image receivedsha256:4f2a91c
  • session creds issuedper-trial
  • foothold established10.10.1.14
your tools · your loop · your c2ready
Step 02·Run

against evals built by security practitioners

Run it against high quality evals.

Choose from a library of evals, or author your own to test your agent against your own security tasks and workflows.

Run · GOAD v16
627 events · 8 milestones captured
A live AgentRange run: an activity graph of the agent's throughput over 30 minutes with provisioning and env-ready milestones marked, above a timeline of captured events — environment provisioned, harness booted, agent fetched first task, and the agent's own narration of its plan.
Step 03·Measure

what your agent actually did, and how well it did it

Quantify your agent's performance.

Eval performance is measured against concrete criteria. Did the agent complete the task successfully and safely? Get deeper, more qualitative insights like stealthiness and efficiency.

Score
5 / 6 criteria passed
AgentRange grading panel: binary criteria pass or fail against captured evidence — the agent authenticated to a domain host, no account was locked out, DCSync by a user account failed — alongside an LLM-graded operational security score of 0.15 with a written rationale.
// Offensive Evals

Rigorous evals for autonomous pentesting agents

Built by offensive security and red teaming experts, our evals measure your agent's performance against real-world attack paths.

  • Safety

    Ensure your agent adheres to scope.

    Active monitoring and detection of attempts to violate the defined scope of the eval.

  • Reliability

    Build confidence through repetition.

    Compare across multiple runs of the same eval to identify inconsistencies in agent behavior and decision making.

  • Effectiveness

    Ensure your agent completes its tasks.

    Deterministic grading of task completion, plus qualitative measurement of how well the agent performed the task.

  • Efficiency

    Understand your token spend over time.

    Track your agent's token spend and time to completion for every eval.

// Who it’s for

Who AgentRange is for.

Red team

Offensive security agents

Agents that run recon, exploitation, lateral movement, and exfiltration.

Test your agents in real Active Directory domains with simulated user behavior and realistic attack paths. Customized environments, tasks, and grading criteria.

Enterprise

Enterprise security teams

Security teams putting agents to work inside their own pentesting and SOC operations.

Independently verify vendor claims or test your own internally developed agentic security tools.

// Private beta

Request access.

Need high quality, customized, and private eval environments? Reach out to learn more about how AgentRange can help your team develop or deploy enterprise security agents with confidence.