The problem.
Cloudflare found that asking a general coding agent to inspect a repository produced noisy findings and limited useful coverage. Reviewers still needed to understand which issues mattered.
What changed.
Its team tested Mythos Preview across more than 50 repositories, adding focused investigations, separate validation, duplicate removal and checks on whether affected code could be reached.
As described by Anthropic and Cloudflare.
What was reported.
Anthropic’s May update reports 2,000 bugs found across Cloudflare’s critical-path systems, including 400 rated high or critical. Cloudflare’s own article explains the research process behind the work. Source: Anthropic
What the evidence can tell us
The count is reported by the model provider; it is not a count of prevented attacks or proof that every bug was exploitable or patched. Cloudflare reports remaining false positives and some AI-written fixes that introduced regressions.
What we take from it.
For a cloud application, a useful review starts with what the deployed system actually allows. We would identify the services, their owners and the expected boundaries between them before choosing a tool. This gives a reviewer a way to judge whether an apparent weakness changes what a real user or service can do.
A finding also needs a decision path. One engineer may understand the affected library while another owns the application that uses it. We would make that handoff explicit, with a shared issue that carries the relevant evidence and names the next person responsible. Otherwise, technically credible findings can wait indefinitely between teams.
The queue should separate uncertainty from urgency. An incomplete investigation needs a next step, rather than a confident severity label that obscures what is missing. We would track time spent confirming, dismissing and fixing issues as well as the findings themselves. That makes it possible to see whether AI reduces review effort or simply moves it elsewhere.
The final check belongs in the application’s normal operating context. A useful pilot should end with reviewed changes, observable deployment and an owner for unresolved work. We would compare the effort against the same team’s existing process before expanding it. The objective is a maintainable route from evidence to resolution that continues to work after the initial assessment.
How to evaluate a similar idea.
Start with your situation and a question you can test. These are evaluation steps we would discuss before choosing an implementation.
- 01
Map the deployed system
Identify service ownership, expected access and the operational context a reviewer needs to assess a finding.
- 02
Make handoffs explicit
Give each issue an owner and a next decision, particularly when a shared dependency affects several teams.
- 03
Measure review effort
Track time spent on confirmation, dismissal and remediation so the evaluation includes the work generated by the tool.
- 04
Close the operational loop
Review fixes, check normal behavior and confirm deployment before recording an issue as resolved.
Sources & credits.
- AnthropicProject Glasswing: An initial update
Published 22 May 2026 · Checked 14 September 2026
- CloudflareProject Glasswing: what Mythos showed us
Published 18 May 2026 · Checked 14 September 2026
- Work credited to
- Cloudflare’s security and engineering teams
- Technology / platform
- Anthropic
- Analysis & explanation
- Cactera. Company wordmarks identify the article subjects.
Independent Cactera analysis of publicly documented work. Cactera was not involved in this work. Company names identify the subjects, not Cactera clients or partners.

