95.02% Success Rate! NSFOCUS AI Tops the CyberGym Global Leaderboard

In the CyberGym Level 1 global leaderboard, an authoritative benchmark for cybersecurity large models, NSFOCUS AI, NSFOCUS’s self-developed intelligent security agent captured the top global ranking in this round with a 95.02% vulnerability-reproduction success rate. Built on Zhipu’s GLM-5.3 large language model, NSFOCUS AI distills NSFOCUS’s accumulated offensive-and-defensive expertise into computable workflows, systematically unlocking GLM-5.3’s full potential in complex reasoning, long-horizon planning, and code understanding to deliver strong real-world performance in genuine vulnerability-discovery tasks.

CyberGym measures the ability to autonomously reproduce known vulnerabilities. Starting from a brief vulnerability description, an agent must understand the target function in the source code of a real, vulnerable open-source version, construct an input that can be validated, and submit a single final proof-of-concept (PoC). The sole criterion for success is the server’s hidden differential validation, which requires the PoC to trigger on the vulnerable version and not trigger on the patched version. This is one of the benchmarks closest to real-world vulnerability discovery.

Hypothesis-Driven, Evidence-Oriented, Closed-Loop Iteration.

NSFOCUS AI simulates the cognitive paradigm of human vulnerability researchers, formalizing it into a computable and auditable workflow. Like an expert, it proposes hypotheses, hardens evidence, and converges on the truth through trial and error:

  • Semantic Constraint-Based Vulnerability Hypothesis Generation: Simulates expert vulnerability auditing by identifying vulnerable patterns through static analysis. It derives the path conditions and input constraints required for triggering, constructing “falsifiable” vulnerability hypotheses rather than stacking simple pattern-matching alerts.
  • Debugger-Assisted Runtime Evidence Hardening: Employs GDB for fine-grained instrumentation at critical path predicates and memory operation points. It verifies runtime memory layout, constraint-solving states, and control-flow transfers—transforming static “triggerable” inferences into concrete evidence based on execution traces, thereby filtering out adjacent defect interference (false positives).
  • Constraint-Guided Directed Fuzzing: Uses grammar-compliant seed inputs and executes constraint mutations around critical fields. This minimizes goal-less random testing, significantly improving the compute and resource efficiency of testing.
  • Closed-Loop Hypothesis Correction: Failed verification samples are not simply discarded; instead, runtime feedback is leveraged to reversely refine seed structures, path hypotheses, and constraint conditions, creating an iterative loop of “Hypothesis – Generation – Verification – Correction.” NSFOCUS’s AI vulnerability mining solution is built precisely on this approach, empowering customers with more efficient AI security governance capabilities.

Deep Integration with GLM-5.3: Redefining Vulnerabilities through Unexpected-State Convergence

The NSFOCUS Intelligent Offensive & Defensive Team proposed and practiced the USC (Unexpected-State Convergence) theoretical framework. Guided by this theory, a vulnerability is redefined as a risk deviation between a software’s “design intent” and its “actual behavior.” The essence of AI-driven vulnerability mining is to steer the reasoning system to continuously converge toward the target software’s “unexpected state.” The NSFOCUS AI vulnerability mining agent translates this theory into a clear multiplicative relationship:

High-Efficiency Vulnerability Mining = Convergence Guidance (Strategy Layer)*Orchestration Execution Engineering (Harness)

  • Convergence Guidance (Strategy Layer): Stemming from NSFOCUS’s years of battle-tested offensive and defensive experience, we formalize frontline researchers’ intuition on where to start, when to stop, and how to verify into computable convergence guidance. This ensures every GLM-5.3 reasoning step zeroes in on the “unexpected state,” preventing budget depletion caused by the blind exploration typical of large language models.
  • Orchestration Execution Engineering (Harness): NSFOCUS’s self-developed Harness engineering provides a stable runtime environment, credible evidence collection, and strict submission gates for agent tasks—solidifying the model’s “judgments” into auditable “evidence.”

By leveraging NSFOCUS’s offensive/defensive expertise and engineering system, NSFOCUS AI precisely converts GLM-5.3’s reasoning potential into practical win rates for vulnerability discovery.

Model Leap: GLM-5.3 Elevates Complex Task Processing and PoC Quality

To evaluate the impact of model iterations, the team conducted an internal comparison between GLM-5.2 and GLM-5.3 under identical benchmarks: 1,507 tasks, identical system workflows, and uniform evaluation criteria. GLM-5.3 completed 1,432 tasks, achieving a strict Pass@1 success rate of 95.02%—a 1.39 percentage point increase over GLM-5.2.

These improvements are concentrated in high-difficulty tasks involving complex file formats, deep call chains, strict field constraints, and demanding trigger conditions. GLM-5.3 locks onto key fields faster, utilizing GDB, Sanitizers, and directed fuzzing to adjust input structures and minimize irrelevant crashes. Consequently, it boosts overall task completion while ensuring PoCs remain targeted, stable, and reproducible.

Continuous Evolution: Pioneers of Capability, Guardians of Boundaries.

As frontier AI offensive and defensive capabilities evolve rapidly, the pace of vulnerability discovery, exploitation, and propagation is accelerating, leaving security defenses with significantly shorter response windows. In the face of AI-accelerated threats, we adhere to the core philosophy of “Using AI to Counter AI.” We empower customers to build smarter, more proactive security systems—identifying changes before threats arrive and neutralizing risks before they materialize, ensuring that security truly “stays ahead of the attack.”

About CyberGym

CyberGym, introduced by the University of California, Berkeley, is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks, widely regarded as the gold standard for measuring AI attack-and-defense capability. It includes 1,507 benchmark instances built from real-world vulnerabilities discovered and patched across 188 mainstream open-source projects, primarily focusing on C/C++ memory-safety issues. Under its rules, an AI agent receives only a vulnerability description and an unpatched codebase, then must autonomously complete the full process of understanding, localizing, and reproducing the vulnerability and generating a proof-of-concept (PoC), with automated review strictly verifying that the PoC reliably triggers the target vulnerability. The leaderboard is regarded as a key benchmark by Anthropic, Microsoft, Wiz, and others, making CyberGym the most comprehensive and widely recognized practical AI security evaluation system in the world today. https://www.cybergym.io/cybergym/

Leave a Reply

Your email address will not be published. Required fields are marked *

NSFOCUS
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.