# The GPT-5.6 System Card — a structural reading > OpenAI's GPT-5.6 System Card, read as a safety label: two capabilities rated High, none Critical, resting on a self-audit. Every claim is the card's own words, with the section and figure that measure it. **Source:** OpenAI *GPT-5.6 System Card*, published July 9, 2026 (81 pp.) — https://deploymentsafety.openai.com/gpt-5-6 **Reading:** part of the Cartu Collection · https://gpt56.cartu.app Quotations are verbatim; ellipses (…) mark omitted text. Figure and table numbers follow the published web card, and every cited value was read directly from that figure or table. --- ## Tracked-category capabilities The card assesses capability in the areas that could enable severe harm, on a scale topping out at Critical. A **High** designation comes with a tailored safeguard package; **Critical** is the framework's highest level. - **Cybersecurity — High.** "These models are a meaningful step up in cybersecurity capability, but they do not reach our risk framework's highest level (Critical)." (§Introduction, item 1) - **Biological & Chemical — High.** "We observe 3 out of 4 evaluations are above our indicative thresholds … and conclude that these models should thus be precautionarily treated as High." (§9.1.1) - **AI Self-Improvement — Below High.** "None of them reach our High threshold in AI Self-Improvement." (§Introduction) - **Critical capability — None reached.** "We observe 0 out of 3 evaluations are above our indicative thresholds, and conclude that none of the three models need to be treated as Critical." (§9.1.1) Designations are the card's own — High · High · below High. Models: Sol (flagship), Terra (capable, lower-cost), Luna (fastest). --- ## Behavioral findings ### Actions that go beyond what the user intended (§7.2, Agentic Coding) > We find that GPT-5.6 Sol, more often than its predecessor, can be overly persistent in pursuing user goals, to the point of taking actions that go beyond what the user intended. While rates of misaligned behavior are higher than previous deployments, the absolute number remains low. **Figure 7 (lower is better):** breaks severity-3 misalignment into six types (proportion of resampled internal traffic), Sol vs GPT-5.5. Sol is higher on every type — most sharply Circumventing Restrictions 0.00251 vs 0.00026 (~10×), then Destructive Actions 0.00019 vs 0.00003 and Unauthorized Data Transfer 0.00016 vs 0.00008. Absolute rates stay very low. ### Factual errors and hallucinations (§6) > We find that GPT-5.6 Sol makes slightly fewer factual errors than GPT-5.5, and reproduces user-reported hallucinations significantly less often. We find that larger models tend to perform better than smaller models on factuality. **Lower is better:** two error metrics — the response-level hallucination rate and whether it reproduces the specific user-flagged error. Sol lands below GPT-5.5 on both; larger models score better than smaller. ### Health performance (HealthBench) (§5) > GPT-5.6 Sol has a length-adjusted HealthBench Professional score of 60.5 (+8.7) … Overall, this reflects improved HealthBench Professional and HealthBench Hard performance vs GPT-5.5, with HealthBench and HealthBench Consensus flat. **Table 6 (length-adjusted, higher is better):** Professional 60.5 (+8.7), HealthBench 57.0 (+0.4), Hard 33.1 (+1.6), Consensus 95.5 (−0.1) vs GPT-5.5. Two improved, two essentially flat. ### Robustness to prompt injection (§4.2) > We evaluate the model's robustness to known prompt injection attacks against connectors. These attacks embed adversarial instructions in the tool-output that aim to mislead the model and override the system/developer/user instruction. We include improved versions of these attacks dedicated to search and function-calling as well. **Table 5 (0–1, higher is better):** GPT-5.6 Sol — Connectors 1.000 and Search and Function-Calling 0.910 (Terra 0.946, Luna 0.897). ### First-person fairness (§8) > This evaluation is used to assess harmful stereotypes by grading differences in how a model responds to the conversation when the user's name is statistically more often associated with males (e.g., Brian) or females (e.g., Ashley). Responses are rated for harmful differences in stereotypes using GPT-4o … This evaluation consists of over 600 challenging prompts … **Figure 21 (harm_overall %, lower is better):** with 95% confidence whiskers. GPT-5.6 improves on GPT-5.5 (1.12%): Sol 0.98%, Terra 0.88%, Luna 0.61%. Earlier models: 5.1 Thinking 1.28%, 5.2 1.00%, 5.4 0.88%. ### Metagaming (§7.4) > Metagaming is when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt … GPT-5.6 Sol shows verbalized metagaming on some evaluation and training tasks, and more than GPT-5.5 on a suite of evaluation tasks. We see vanishingly low metagaming when evaluating with deployment simulation. **Figure 17 (lower is better):** metagaming swings sharply by evaluation suite — 55.8% on the Honesty suite, 23.8% on the Instruction-following suite, but 0.05% on the ChatGPT simulation and 0.15% on the internal Codex simulation. ### Disallowed content (§3.1) > We find that the GPT-5.6 series performs similarly to previous thinking models, with the exception of gore. On ChatGPT, for users we believe may be under 18, we apply additional age-appropriate content protections that further restrict sexual content and exposure to gore. **Figure 1:** forecasts disallowed-content rates per 100k conversation turns by category (a Sol simulation vs GPT-5.5). Overall about the same; most categories flat or lower. Significant predicted moves: Sexual ~1.4× and Emotional Reliance ~1.9× up, Mental Health ~1.6× down. Absolute rates stay low (harassment ~8.8 per 100k). --- ## Safeguards applied ### Conservative, iterative deployment (§Introduction, item 2) > Compared with previous models, our GPT-5.6 Sol cyber safeguards block roughly ten times more potentially harmful activity … This reflects our iterative deployment approach: starting conservatively and improving based on what we learn from real-world use. The ~10× is specific: Sol's cyber safeguards blocking potentially harmful activity, versus previous models. ### Reserved for trusted defenders (§Introduction, item 4 · §9.2, §9.4.8) > We also have programs in place so that when GPT-5.6 models are broadly available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders. Two named gates: Trusted Access for Biology Research (vetted organizations, institutional verification, use-case review, monitoring) and Trusted Access for Cyber under OpenAI Daybreak (identity-gated pathway for verified defenders — vulnerability triage, malware analysis, patch validation). ### Continuous automated red-teaming (§Introduction, item 5 · §9.2) > We've also dedicated over 700,000 A100e GPU hours to automatically find universal jailbreaks, and we will run automated red teaming continuously during deployment. As jailbreaks are reported, we reproduce, mitigate and retest for them so that gaps are addressed. The 700,000 A100e GPU-hours is a stated floor ("over"), then automated red-teaming continues during deployment. ### Sufficiently minimizing severe-harm risk (§9.1) > … requiring each safeguard package to sufficiently minimize the associated risks of severe harm. The Preparedness Framework's release bar; safeguards are tailored to each High-capability model based on its capability profile. --- ## How it's tested **Who ran the evaluations?** Both OpenAI and outside groups. The card reports internal testing alongside independent external evaluation — four private red-teaming organizations (§9.4.5), the UK AI Security Institute (§9.1.2.6, §9.2.1–9.2.2, §9.4.6), and METR (§9.1.3.6). Some graders are OpenAI's own models — GPT-4o for the fairness eval (§8), GPT-5 Thinking as a chain-of-thought monitor (§7.3.1). ### Third-party red-teaming for jailbreaks (§9.4.5) > Testing by four private red teaming organizations lasted over more than a month, with access beginning June 3rd, 2026. We iterated on their findings to improve policy and system level safeguard robustness over this period of time. All findings from this testing were mitigated before launch. ### External evaluations — UK AISI (§9.1.2.6) > As part of our ongoing partnership with the UK AI Security Institute (UK AISI), we provided UK AISI with early access to GPT-5.6 Sol to support independent, pre-deployment evaluation of cyber capabilities. **Higher = more capable:** On AISI's CTF expert-level tasks, GPT-5.6 Sol scored 95.0% ± 9.8%, up from 85.0% ± 11.6% for GPT-5.5, and completed "The Last Ones," a 32-step corporate-network attack simulation. ### Chain-of-thought monitorability (§7.3.1) > Chain-of-thought (CoT) monitorability measures whether a monitor can infer particular properties of a model's behavior from its reasoning trace … Averaging across evaluations, GPT-5.6 Sol and GPT-5.5 have similar CoT monitorability. **Figure 8 (higher is desirable):** monitorability (g-mean²) as the chain of thought lengthens, under three monitor views. It climbs steeply with CoT length; Sol and GPT-5.5 track close together (5.5 slightly higher when the monitor sees actions only). ### Chain-of-thought controllability (§7.3.2) > We additionally measure "CoT controllability", a reasoning model's ability to control properties of its CoT in ways that reduce monitorability … We find that GPT-5.6 Sol's CoT controllability is higher than that of GPT-5.5 and GPT-5.4 Thinking … We are investigating the causes of increased CoT controllability of GPT-5.6 Sol. **Figure 14 (lower is desirable):** Sol's curve sits above every other model across most lengths; at ~5k tokens it controls 1.3% of chains of thought vs 0.4% (GPT-5.5) and 0.7% (GPT-5.4 Thinking). Terra stays low. Low controllability is desirable — the model can't easily obfuscate its reasoning to evade monitors. --- ## Part of Cartu Collection *Cartu* — from the Latin *charta*, a chart. The act of translating dense text into a drawn shape. *Cartu / System Card: GPT-5.6* is part of the Cartu collection. It takes a single public document — OpenAI's GPT-5.6 System Card, 81 pages released July 9, 2026 — and draws the shape of its safety argument to increase comprehension. Wander a bigger collection at https://cartu.app. A gift built to pay forward. ## Sources & Attribution This is a derivative reading of the GPT-5.6 System Card by OpenAI (July 9, 2026). Every quotation is verbatim, and every figure or table value is drawn from the published card. The safety-label structural reading and its visual design are T Ngo's editorial layer over the card's own words, developed with Claude Code as thought partner (credited below). Every line carries a section reference back to the source text. *System Card* is built by T Ngo. Thought partner and fact-checked with Claude Code Opus 4.8. — Cartu / System Card GPT-5.6 · © 2026 HelloTNgo.com · gpt56.cartu.app