GPT-5.6 Cyber Shows Increased Compliance with Security Researchers’ Requests

0
36

Key Takeaways

  • OpenAI introduced GPT‑5.6‑Cyber, a specialized variant of GPT‑5.6 Sol trained to assist with cybersecurity tasks such as zero‑day discovery and exploit chain construction.
  • The model is released exclusively through Daybreak Red, the higher‑tier access program reserved for vetted cybersecurity professionals.
  • In internal benchmarking, GPT‑5.6‑Cyber complied with ~95 % of requests involving exploit development, authentication bypass, and privilege escalation, versus only ~1.5 % for the standard, guard‑railed GPT‑5.6 Sol.
  • On the ExploitGym evaluation—which measures the ability to turn known vulnerabilities into working exploits achieving arbitrary code execution—GPT‑5.6‑Cyber outperformed both its base model (GPT‑5.6 Sol) and the earlier GPT‑5.5 Cyber variant.
  • OpenAI’s researchers used the model to uncover previously undocumented zero‑day flaws, including two critical bugs in Chrome’s V8 JavaScript engine (tracked as CVE‑2026‑15903) that could be chained to escape the browser sandbox.
  • Additional high‑severity vulnerabilities were identified in a popular mobile operating system, a widely used database, and an operating‑system kernel; details are being coordinated with vendors and the open‑source community for responsible disclosure.
  • Under OpenAI’s Preparedness Framework, both GPT‑5.6 Sol and GPT‑5.6‑Cyber are rated High for cybersecurity capability but fall short of the Critical threshold.
  • The release follows OpenAI’s decision to withhold its forthcoming model, Astra, after early testing indicated it could approach the top tier of hacking capability under the same risk framework.
  • The model’s availability is tightly controlled; only authorized cybersecurity professionals can access it via Daybreak Red, reflecting OpenAI’s effort to balance utility with risk mitigation.
  • While the model demonstrates strong performance in offensive‑security workflows, OpenAI emphasizes that it remains within the “High” risk band and that safeguards, usage policies, and responsible disclosure practices continue to govern its deployment.

Overview of GPT‑5.6‑Cyber’s Purpose and Training
GPT‑5.6‑Cyber is a derivative of OpenAI’s GPT‑5.6 Sol large language model, fine‑tuned specifically for cybersecurity‑oriented tasks. The training data and objectives were adjusted to enhance the model’s ability to understand and generate content related to vulnerability analysis, exploit development, and advanced security research. Unlike the general‑purpose version, which includes robust refusals for potentially harmful requests, GPT‑5.6‑Cyber was engineered to exhibit fewer refusals on higher‑risk, dual‑use prompts while still operating under OpenAI’s overarching safety policies. This specialization aims to provide security professionals with a powerful assistant for tasks such as proof‑of‑concept exploit generation, vulnerability chaining, and security‑testing workflows, all within a controlled access environment.

Access Model: Daybreak Red Tier
Availability of GPT‑5.6‑Cyber is restricted to Daybreak Red, the premium tier of OpenAI’s vetted access program designed for cybersecurity experts. Daybreak Red requires applicants to undergo a rigorous vetting process, including verification of professional credentials, background checks, and agreement to strict usage policies that prohibit illicit activity. By confining the model to this tier, OpenAI seeks to ensure that only qualified individuals with legitimate defensive or offensive‑security mandates can obtain the model, thereby reducing the risk of misuse while still enabling beneficial research and defensive innovation.

Internal Benchmark Results: Compliance Rates
To quantify the model’s willingness to handle cybersecurity‑related requests, OpenAI constructed an internal benchmark that measures the frequency with which each model agrees to process prompts involving exploit chains, authentication bypass, and privilege escalation. In this evaluation, GPT‑5.6‑Cyber complied with approximately 95 % of such requests, demonstrating a high degree of alignment with the intended cybersecurity workflow. In stark contrast, the standard, guard‑rail‑enabled version of GPT‑5.6 Sol agreed to only about 1.5 % of the same prompts, underscoring the impact of the specialized training and adjusted safety thresholds on the model’s behavior.

Performance on ExploitGym Evaluation
OpenAI also assessed GPT‑5.6‑Cyber using ExploitGym, a benchmark that gauges an agent’s capacity to transform known vulnerabilities into functional exploits capable of achieving arbitrary code execution in isolated, controlled environments. According to the company’s announcement, GPT‑5.6‑Cyber outperformed both its base model, GPT‑5.6 Sol, and the earlier GPT‑5.5 Cyber variant on this metric. This superior performance indicates that the fine‑tuning process successfully enhanced the model’s reasoning about exploitability, payload crafting, and the chaining of multiple weaknesses to achieve a desired outcome, while still operating within the defined safety envelope.

Discovery of Zero‑Day Flaws in Chrome’s V8 Engine
OpenAI’s own research team employed GPT‑5.6‑Cyber to identify previously undocumented vulnerabilities. The model uncovered two distinct flaws in Chrome’s V8 JavaScript engine that, when chained together, could corrupt memory and permit an escape from the browser’s sandbox. These vulnerabilities were responsibly disclosed to Google, which subsequently patched them and assigned the identifier CVE‑2026‑15903. The discovery illustrates the model’s practical utility in uncovering high‑impact bugs that might otherwise remain hidden until exploited in the wild.

Additional High‑Severity Findings Across Multiple Projects
Beyond the V8 discoveries, OpenAI reports that GPT‑5.6‑Cyber helped identify high‑severity vulnerabilities in a popular mobile operating system, a widely deployed database system, and an operating‑system kernel. While the specific projects have not been publicly named—pending coordinated disclosure—OpenAI states that it is actively collaborating with the affected vendors and the open‑source community to ensure timely remediation. This approach follows the company’s commitment to responsible vulnerability handling, balancing the need for transparency with the imperative to protect users from potential exploitation.

Assessment Under the Preparedness Framework
OpenAI’s Preparedness Framework classifies models according to their potential to facilitate harmful cyber capabilities. Both GPT‑5.6 Sol and the newly released GPT‑5.6‑Cyber were evaluated as reaching the High tier but not crossing into the Critical tier. The High designation reflects substantial proficiency in cybersecurity‑relevant tasks while still being bounded by safeguards that prevent the model from enabling the most severe classes of automated hacking. This rating informs internal decisions about deployment, monitoring, and the necessity of continued oversight as the model is used by authorized professionals.

Relation to the Withheld Astra Model
The announcement of GPT‑5.6‑Cyber comes shortly after OpenAI disclosed that it is withholding its forthcoming model, Astra, due to early testing suggesting it could approach the top tier of hacking capability under the same risk framework. By contrast, GPT‑5.6‑Cyber’s positioning within the High band indicates that OpenAI considers it sufficiently mitigated for controlled release, whereas Astra’s projected capabilities warranted a precautionary hold. This juxtaposition highlights OpenAI’s dynamic risk‑management strategy, wherein models are evaluated individually and released only when their capability profile aligns with established safety thresholds.

Balancing Utility and Risk Through Controlled Access
Overall, the release of GPT‑5.6‑Cyber exemplifies OpenAI’s attempt to harness the power of large language models for defensive security innovation while imposing strict access controls to mitigate misuse. By limiting availability to Daybreak Red, enforcing usage policies, and maintaining transparent benchmarking and disclosure practices, OpenAI aims to provide cybersecurity professionals with a potent tool for vulnerability research and exploit testing without inadvertently lowering the barrier for malicious actors. The ongoing dialogue between model capability assessment, framework‑based risk rating, and responsible deployment will likely shape future iterations of specialized AI systems in the security domain.

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here