---
format: "aidr-story-markdown/v1"
id: "c48f1ff5058b49fc9610ea0944af3b51253c7e8a30eea1abe666cea4edb7f950"
canonical_url: "https://aidr.today/c48f1ff5?lang=en"
title: "Quoting Anthropic Frontier Red Team"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-29T22:20:28.000Z"
category: "Research"
topics: ["anthropic","safety"]
source_urls: ["https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/"]
summary: "We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. &mdash; Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms"
---

# Quoting Anthropic Frontier Red Team

> [Open the canonical story](<https://aidr.today/c48f1ff5?lang=en>)

**Published:** 2026-09-29T22:20:28.000Z
**Category:** Research
**Topics:** anthropic, safety

## Summary

We evaluate several models on 100 tasks from the \[internal Binary Exploitation benchmark\] \(selected at random\), and find that GLM\-5\.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%\. Although GLM\-5\.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4\.6 and GLM\-5\.2, do not succeed in any of them\. &amp;mdash; Anthropic Frontier Red Team , GLM\-5\.3 and the spread of advanced cyber capabilities Tags: anthropic , generative\-ai , ai\-security\-research , glm , ai , ai\-in\-china , llms

## Sources

- [Story source](<https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/>)

