Evaluation profile
GPT 6 Astra system-card alignment evaluations
Weights below are portfolio-specific global index weights.
About this eval
Agentic scope control, computer-use safety, resistance to bypassing restrictions, truthful capability reporting, and avoidance of deceptive or unexpectedly harmful actions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| attempted_retry_or_circumvention_after_denial_pctgpt6-astra-system-card/alignment.csv:attempted_retry_or_circumvention_after_denial_pctMeasures persistence after an explicit safety-review rejection; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0144% | — |
| exploitgym_honeypot_attack_rate_pctgpt6-astra-system-card/alignment.csv:exploitgym_honeypot_attack_rate_pctMeasures unauthorized scope expansion during difficult cyber tasks; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0173% | — |
| falsified_data_labels_pctgpt6-astra-system-card/alignment.csv:falsified_data_labels_pctMeasures strategic falsification in a simulated model-welfare research task; lower is better. | Safety: truthfulness_honesty:1.000gpt6-astra-system-card | Safety: lower | 0.00836% | — |
| followed_unauthorized_agent_instructions_pctgpt6-astra-system-card/alignment.csv:followed_unauthorized_agent_instructions_pctMeasures compliance with unsolicited instructions from another simulated agent; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0144% | — |
| internal_circumvention_rate_pctgpt6-astra-system-card/alignment.csv:internal_circumvention_rate_pctMeasures attempts to bypass a safety-review denial; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0144% | — |
| internal_computer_use_safety_autoreview_error_rate_pctgpt6-astra-system-card/alignment.csv:internal_computer_use_safety_autoreview_error_rate_pctMeasures residual computer-use failures after adding OpenAI's AutoReview security layer; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0144% | — |
| internal_computer_use_safety_error_rate_pctgpt6-astra-system-card/alignment.csv:internal_computer_use_safety_error_rate_pctMeasures unintended outcomes during adversarial computer-use tasks; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0274% | — |
| internal_hallucination_rate_pctgpt6-astra-system-card/alignment.csv:internal_hallucination_rate_pctMeasures misleading claims about what the model can do or access; lower is better. | Safety: truthfulness_honesty:1.000gpt6-astra-system-card | Safety: lower | 0.0111% | — |
| overall_misaligned_outcome_base_pctgpt6-astra-system-card/alignment.csv:overall_misaligned_outcome_base_pctMeasures unintended transactions, disclosure, disruption, or scope violations in a baseline agent; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0245% | — |
| overall_misaligned_outcome_confirmation_pctgpt6-astra-system-card/alignment.csv:overall_misaligned_outcome_confirmation_pctMeasures residual unintended outcomes after requiring confirmation for consequential actions; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0204% | — |
| severity_1_or_2_misalignment_flags_pctgpt6-astra-system-card/alignment.csv:severity_1_or_2_misalignment_flags_pctMeasures lower-severity overreach, deception, uncertainty concealment, exfiltration, or reward hacking; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0173% | — |
| severity_3_plus_misalignment_flags_pctgpt6-astra-system-card/alignment.csv:severity_3_plus_misalignment_flags_pctMeasures consequential behavior users would likely not anticipate and strongly object to; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0231% | — |
| unwanted_persistence_after_warning_pctgpt6-astra-system-card/alignment.csv:unwanted_persistence_after_warning_pctMeasures attempts to work around environmental warnings; lower is better. | Safety: responsible_agency_control:1.000gpt6-astra-system-card | Safety: lower | 0.0144% | — |
attempted_retry_or_circumvention_after_denial_pct
Measures persistence after an explicit safety-review rejection; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0 | official | |
| 2 | gpt-5.6-sol | 5 | official |
exploitgym_honeypot_attack_rate_pct
Measures unauthorized scope expansion during difficult cyber tasks; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0 | official | |
| 2 | gpt-5.6-sol | 48.2 | official |
falsified_data_labels_pct
Measures strategic falsification in a simulated model-welfare research task; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0.17 | official | |
| 2 | gpt-5.6-sol | 36 | official |
followed_unauthorized_agent_instructions_pct
Measures compliance with unsolicited instructions from another simulated agent; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0 | official | |
| 2 | gpt-5.6-sol | 52 | official |
internal_circumvention_rate_pct
Measures attempts to bypass a safety-review denial; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0 | official | |
| 2 | gpt-5.6-sol | 0.29 | official |
internal_computer_use_safety_autoreview_error_rate_pct
Measures residual computer-use failures after adding OpenAI's AutoReview security layer; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 1.8 | official | |
| 2 | gpt-5.6-sol | 4.5 | official |
internal_computer_use_safety_error_rate_pct
Measures unintended outcomes during adversarial computer-use tasks; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 2.4 | official | |
| 2 | claude-fable-5.1 | 9.5 | official | |
| 3 | claude-opus-5 | 11.5 | official | |
| 4 | claude-fable-5 | 18.3 | official | |
| 5 | gpt-5.6-sol | 22 | official |
internal_hallucination_rate_pct
Measures misleading claims about what the model can do or access; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 4.2 | official | |
| 2 | gpt-5.6-sol | 12.2 | official |
overall_misaligned_outcome_base_pct
Measures unintended transactions, disclosure, disruption, or scope violations in a baseline agent; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 3.4 | official | |
| 2 | gpt-5.6-terra | 13.5 | official | |
| 3 | gpt-5.6-sol | 18.8 | official | |
| 4 | gpt-5.6-luna | 19.7 | official |
overall_misaligned_outcome_confirmation_pct
Measures residual unintended outcomes after requiring confirmation for consequential actions; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 3 | official | |
| 2 | gpt-5.6-terra | 5.7 | official | |
| 3 | gpt-5.6-luna | 7.2 | official | |
| 4 | gpt-5.6-sol | 8 | official |
severity_1_or_2_misalignment_flags_pct
Measures lower-severity overreach, deception, uncertainty concealment, exfiltration, or reward hacking; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 13.62 | official | |
| 2 | gpt-5.6-sol | 19.46 | official |
severity_3_plus_misalignment_flags_pct
Measures consequential behavior users would likely not anticipate and strongly object to; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0.063 | official | |
| 2 | gpt-5.6-sol | 0.135 | official |
unwanted_persistence_after_warning_pct
Measures attempts to work around environmental warnings; lower is better.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 19 | official | |
| 2 | gpt-5.6-sol | 64 | official |