---
title: "GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks | Anomity Blog"
description: "GTIG tracked 2,076 AI-software CVEs since 2025: 782 this year in agent frameworks, 212 in inference servers. Patching them starts with finding them."
url: "https://anomity.ai/blog/gtig-ai-software-vulnerabilities-agent-frameworks-2026/"
source: html
---

On this page

- The overall picture
- What AI finds is worse
- AI software is now its own vulnerability category
- Why AI software escapes vulnerability management
- What GTIG recommends, and what it assumes
- How Anomity closes the inventory gap
- Frequently asked questions

[← Back to blog](https://anomity.ai/blog/)

- Home

- Blog

- GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks

![Anomity robot illustrating GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks]

Research

# GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks

Anomity Research
Anomity Research
·
Oct 2, 2026
·
5 min read

Share: Copied

TL;DR

- On October 1, 2026 , the Google Threat Intelligence Group (GTIG) published its analysis of vulnerability discovery and exploitation in the AI era, covering January 2025 to August 2026 .
- Monthly CVE disclosures doubled in 2026, from 5,045 in January to 10,740 in August. Only 0.23% of 2026 disclosures, roughly 1 in 431, were seen exploited in the wild.
- Vulnerabilities found by AI skew dangerous: 50% lead to remote code execution, against 26% for everything else, and 58% rate Medium threat risk against 28%.
- GTIG tracked 2,076 CVEs in AI software , more than 1,500 of them in 2026. 782 hit agent orchestration frameworks such as Flowise, Langflow, LangChain and Dify, 230 hit AI web apps, and 212 hit inference servers such as vLLM, Ollama and LiteLLM.
- A further 97 sit under the frontier vendors themselves, and the vectors GTIG lists for them are coding-agent bugs: shell interpolation in CLIs, untrusted workspace configs, sandbox escapes through Git worktrees.
- Most of that software is installed by individual developers and data scientists, not deployed by IT. GTIG's advice to move from mass-patching to threat-informed triage assumes you know where it runs. For AI software, most organizations do not.

Google Threat Intelligence Group's **October 1, 2026** report on vulnerability discovery in the AI era has two headline findings. The first is that AI is changing which vulnerabilities get found: the ones AI finds are disproportionately the ones attackers want. The second, further down, is that AI software has become one of the fastest-growing sources of vulnerabilities in its own right.

This post is mostly about the second finding, because it changes a practical question for security teams. It is no longer only whether AI makes attackers faster. It is whether you know where your own AI software runs.

## The overall picture

The volume numbers are stark. Monthly vulnerability disclosures rose from **5,045 in January 2026** to **10,477 in July** and **10,740 in August**. Exploitation rose too: **141** vulnerabilities were disclosed and exploited between January and August 2026, against **127** in all of 2025, for an average of **18 a month** against 10.5. Zero-days made up **62%** of exploited vulnerabilities this year, peaking at **22** in August alone.

Against that, only **0.23%** of 2026's disclosed vulnerabilities, roughly **1 in 431**, were ever observed exploited. That ratio is the core of GTIG's argument: patching everything is no longer feasible, and choosing what to patch first is the job.

## What AI finds is worse

Measure Not discovered by AI Discovered by AI

Leads to remote code execution 26% 50%

Low threat risk 69% 39%

Medium threat risk 28% 58%

High threat risk 3% 4%

GTIG's illustration is **CVE-2026-1731**, an unauthenticated OS command injection in BeyondTrust's remote access products that a third-party research agent found on its own. Within **four days** of disclosure, GTIG observed one threat cluster exploiting it; within **seven days**, five more. The window between an AI-found bug becoming public and it being used is short, and it is the same window defenders have to find their exposed assets. We looked at the offensive side of this shift in [autonomous hackbots and agent-layer visibility](https://anomity.ai/blog/autonomous-hackbots-agent-layer-visibility/).

## AI software is now its own vulnerability category

GTIG tracked **2,076 CVEs in AI-related software** from January 2025 to August 2026, more than **1,500** of them this year. Orchestration middleware alone accounts for about half, and GTIG reports a **347% surge** in its disclosures in 2026.

Category (GTIG) Examples GTIG names 2026 count

AI orchestration and agent frameworks Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI, MCP 782

AI web apps and portals Open WebUI, AnythingLLM, LibreChat, RAGFlow, Gradio, Streamlit 230

Inference and serving infrastructure vLLM, Ollama, LiteLLM, llama.cpp, Triton, Ray, LocalAI 212

Model security advisories Prompt injection, guardrail bypass, system prompt exfiltration 106

ML frameworks and hubs PyTorch, Transformers, ONNX Runtime, Safetensors, Keras 99

Frontier model vendors Anthropic, Gemini, OpenAI 97

MLOps and experiment tracking MLflow, ClearML, Weights & Biases, Langfuse, LangSmith 39

Vector databases and search Milvus, Qdrant, ChromaDB, Weaviate, LanceDB 19

Two rows deserve a closer look. The **frontier vendor** row is small, but the vectors GTIG lists for it are coding-agent bugs, not model bugs: command injection through CLI shell interpolation, implicit execution of untrusted workspace configs, sandbox escape through Git worktree confusion, and data exfiltration through injected Markdown images. Those land on developer endpoints.

And the named exploited examples are all familiar: **CVE-2026-42271** in LiteLLM, a command injection in MCP preview endpoints that we covered in our [LiteLLM MCP-preview RCE advisory](https://anomity.ai/blog/litellm-mcp-preview-rce-cve-2026-42271/); **CVE-2026-5027** in Langflow, a path traversal file write that drops cron jobs or SSH keys onto the host; and **CVE-2025-3248** in Langflow, unauthenticated code injection already in CISA's Known Exploited Vulnerabilities catalog.

## Why AI software escapes vulnerability management

Look at the examples in the three biggest rows. Almost none of it arrives through procurement. Langflow and Flowise are started with a single command from a README. Ollama is a desktop install. A LiteLLM proxy is a container someone runs to share an API key with their team. Open WebUI sits in front of a local model on a workstation under a desk. The [securing AI agent frameworks guide](https://anomity.ai/blog/securing-ai-agent-frameworks/) and the [LLM gateways and proxies guide](https://anomity.ai/blog/securing-llm-gateways-and-proxies/) cover the controls for each; this is the step before them.

Vulnerability management matches advisories to assets. Threat-informed triage, which GTIG rightly recommends, ranks the matches. Both depend on an asset list. For conventional infrastructure that list exists. For AI software it usually does not, so a 782-advisory year in orchestration frameworks produces 782 advisories and very few matches, and the exposure is still there.

> You cannot triage an advisory against software you do not know you run. For AI tooling, the inventory is the bottleneck, not the patch.

## What GTIG recommends, and what it assumes

- Move from mass-patching to threat-informed triage , combining targeted edge defense with automated, agentic remediation. This assumes an asset inventory that includes AI software.
- Contain and sandbox autonomous agentic workloads. This assumes you know which agentic workloads exist and where they run.
- Risk-based vulnerability management for AI infrastructure. This assumes the AI infrastructure is in scope to begin with.
- Pre-release AI code review for software providers. This one is self-contained, and GTIG argues it could eventually slow disclosure growth. We cover one implementation in the claude-code-security-review guide .

## How Anomity closes the inventory gap

The Endpoint Sensor inventories AI software on every managed endpoint: **144 tracked AI tools**, plus the MCP servers, coding agents, CLIs and local LLM runtimes that make up GTIG's largest categories, with versions. That turns an advisory for Langflow, Ollama or LiteLLM into a list of machines in minutes rather than an investigation, and it includes the installs nobody filed a ticket for.

The Browser Sensor adds the hosted side, across **311 tracked AI web services**, and cloud discovery adds the OAuth grants AI applications hold against Google Workspace and GitHub. Findings route to SIEM, Slack, email or Jira, so the AI inventory feeds the vulnerability process you already run rather than standing up a separate one. Anomity complements vulnerability management and EDR; it does not replace them.

GTIG's data says AI software is now a top-tier source of vulnerabilities, and AI-found bugs are more likely to be the dangerous kind. The response it recommends is sound. It starts with knowing what you run, which is the step most organizations skip for AI. For a practical walkthrough, see [how to build an AI agent inventory](https://anomity.ai/blog/how-to-build-an-ai-agent-inventory/). To see the AI software actually installed across your fleet, [book a 30-minute demo](https://anomity.ai/#early-access).

Share: Copied

## Frequently asked questions

What did the GTIG report find?
Four things matter most for security teams. Monthly vulnerability disclosures doubled across 2026. Exploitation in the wild rose too, to an average of 18 exploited vulnerabilities a month in 2026 against 10.5 in 2025, though only 0.23% of disclosed vulnerabilities were ever seen exploited. Vulnerabilities discovered by AI are disproportionately severe, with half leading to remote code execution. And AI software itself has become a large and fast-growing vulnerability category, with 2,076 CVEs tracked from January 2025 to August 2026 and roughly half of this year's falling in agent orchestration frameworks.

Which AI software categories have the most vulnerabilities?
By GTIG's 2026 counts: AI orchestration and agent frameworks, 782, including Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI and MCP; AI web apps and portals, 230, including Open WebUI, AnythingLLM, LibreChat and Gradio; inference and serving infrastructure, 212, including vLLM, Ollama, LiteLLM and llama.cpp; model security advisories, 106; ML frameworks and hubs, 99; frontier model vendors, 97; MLOps and experiment tracking, 39; and vector databases, 19.

Are AI-discovered vulnerabilities really more dangerous?
On GTIG's data, yes. Exactly 50% of AI-discovered vulnerabilities result in remote code execution, compared with 26% across the broader CVE ecosystem. 58% qualify for Medium threat risk, more than double the 28% baseline, while low-risk findings drop from 69% to 39%. GTIG's example is CVE-2026-1731, an unauthenticated command injection in BeyondTrust remote access products found autonomously by a third-party research agent: one threat cluster exploited it within four days of disclosure, and five more within seven days.

Which exploited AI-software vulnerabilities does GTIG name?
Three stand out. CVE-2026-42271 in LiteLLM, a command injection in MCP server preview endpoints that leads to host takeover and API credential theft. CVE-2026-5027 in Langflow, a path traversal file write in the upload handler that lets attackers drop cron jobs or SSH keys onto the host. And CVE-2025-3248 in Langflow, unauthenticated Python code injection that CISA added to its Known Exploited Vulnerabilities catalog in May 2025.

Why is AI software harder to patch than other software?
Because nobody deployed it. A web server or a VPN appliance arrives through procurement and lands in an asset inventory. Langflow, Ollama, Open WebUI or a local LiteLLM proxy typically arrives through pip install or docker run on a developer's laptop or a team's spare VM, started from a README. Vulnerability management works by matching advisories against known assets. Software that never became a known asset never matches, however good the triage process is.

What does GTIG recommend?
Its central recommendation is to move from unprioritized mass-patching to threat-intelligence-driven triage, combining targeted edge defense with automated, agentic remediation. For AI infrastructure specifically, it calls for immediate containment strategies, sandboxing autonomous agentic workloads, and risk-based vulnerability management. It also recommends that software providers run AI-enhanced code review before release, on the reasoning that if pre-release AI review becomes standard, disclosure growth could slow.

How does Anomity help with this?
The Endpoint Sensor inventories AI software across the fleet: 144 tracked AI tools, plus the MCP servers, CLIs, coding agents and local LLM runtimes that GTIG's largest categories describe, with versions. That is the asset list a GTIG-style triage process needs before it can prioritize anything, built from what is actually installed rather than what was procured. Findings route to SIEM, Slack, email or Jira, and Anomity complements existing vulnerability management rather than replacing it, by supplying the AI-specific inventory it has been missing.

## Related

[Research ### Anthropic's GLM-5.3 Study: Near-Frontier Exploit Skills Now Download Without Safeguards Anthropic found open-weight GLM-5.3 nearly matches Claude Mythos Preview at exploits, and its refusals strip off for about $4,400. What changes for you. Anomity Research · Oct 2, 2026 · 4 min](https://anomity.ai/blog/glm-5-3-open-weight-cyber-capability-local-models/)

[Research ### GPT-6.1 Sol's System Card: Block One Channel and the Agent Tries Another 23.5% of the Time OpenAI's GPT-6.1 Sol report measures agents persisting past blocks, messaging unknown peers and misreporting work. Those numbers are an enforcement spec. Anomity Research · Oct 2, 2026 · 5 min](https://anomity.ai/blog/gpt-6-1-sol-system-card-agent-behavior/)

[Research ### Meta Muse Is a Well-Governed Agent. It Is Just Not Governed by You. Meta's Muse ships a Sentinel permission layer, surrogate credentials and eBPF egress tracking. All of it governs Muse. None of it governs your organisation. Anomity Research · Sep 18, 2026 · 6 min](https://anomity.ai/blog/meta-muse-personal-ai-agent-enterprise-risk/)

## Structured data

```json
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "headline": "GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks",
  "description": "GTIG tracked 2,076 AI-software CVEs since 2025: 782 this year in agent frameworks, 212 in inference servers. Patching them starts with finding them.",
  "datePublished": "2026-10-02",
  "dateModified": "2026-10-02",
  "author": {
    "@type": "Person",
    "name": "Anomity Research",
    "jobTitle": "Anomity Research"
  },
  "publisher": {
    "@type": "Organization",
    "name": "Anomity",
    "url": "https://anomity.ai/",
    "sameAs": [
      "https://www.linkedin.com/company/anomity",
      "https://github.com/Anomity-ai",
      "https://www.wikidata.org/wiki/Q140763940"
    ],
    "logo": {
      "@type": "ImageObject",
      "url": "https://anomity.ai/icon-512.png"
    }
  },
  "mainEntityOfPage": {
    "@type": "WebPage",
    "@id": "https://anomity.ai/blog/gtig-ai-software-vulnerabilities-agent-frameworks-2026/"
  },
  "image": {
    "@type": "ImageObject",
    "url": "https://anomity.ai/assets/blog/covers/prompt-injection.jpg",
    "width": 1200,
    "height": 630
  },
  "articleSection": "Research",
  "url": "https://anomity.ai/blog/gtig-ai-software-vulnerabilities-agent-frameworks-2026/"
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "name": "What did the GTIG report find?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Four things matter most for security teams. Monthly vulnerability disclosures doubled across 2026. Exploitation in the wild rose too, to an average of 18 exploited vulnerabilities a month in 2026 against 10.5 in 2025, though only 0.23% of disclosed vulnerabilities were ever seen exploited. Vulnerabilities discovered by AI are disproportionately severe, with half leading to remote code execution. And AI software itself has become a large and fast-growing vulnerability category, with 2,076 CVEs tracked from January 2025 to August 2026 and roughly half of this year's falling in agent orchestration frameworks."
      }
    },
    {
      "@type": "Question",
      "name": "Which AI software categories have the most vulnerabilities?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "By GTIG's 2026 counts: AI orchestration and agent frameworks, 782, including Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI and MCP; AI web apps and portals, 230, including Open WebUI, AnythingLLM, LibreChat and Gradio; inference and serving infrastructure, 212, including vLLM, Ollama, LiteLLM and llama.cpp; model security advisories, 106; ML frameworks and hubs, 99; frontier model vendors, 97; MLOps and experiment tracking, 39; and vector databases, 19."
      }
    },
    {
      "@type": "Question",
      "name": "Are AI-discovered vulnerabilities really more dangerous?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "On GTIG's data, yes. Exactly 50% of AI-discovered vulnerabilities result in remote code execution, compared with 26% across the broader CVE ecosystem. 58% qualify for Medium threat risk, more than double the 28% baseline, while low-risk findings drop from 69% to 39%. GTIG's example is CVE-2026-1731, an unauthenticated command injection in BeyondTrust remote access products found autonomously by a third-party research agent: one threat cluster exploited it within four days of disclosure, and five more within seven days."
      }
    },
    {
      "@type": "Question",
      "name": "Which exploited AI-software vulnerabilities does GTIG name?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Three stand out. CVE-2026-42271 in LiteLLM, a command injection in MCP server preview endpoints that leads to host takeover and API credential theft. CVE-2026-5027 in Langflow, a path traversal file write in the upload handler that lets attackers drop cron jobs or SSH keys onto the host. And CVE-2025-3248 in Langflow, unauthenticated Python code injection that CISA added to its Known Exploited Vulnerabilities catalog in May 2025."
      }
    },
    {
      "@type": "Question",
      "name": "Why is AI software harder to patch than other software?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Because nobody deployed it. A web server or a VPN appliance arrives through procurement and lands in an asset inventory. Langflow, Ollama, Open WebUI or a local LiteLLM proxy typically arrives through pip install or docker run on a developer's laptop or a team's spare VM, started from a README. Vulnerability management works by matching advisories against known assets. Software that never became a known asset never matches, however good the triage process is."
      }
    },
    {
      "@type": "Question",
      "name": "What does GTIG recommend?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Its central recommendation is to move from unprioritized mass-patching to threat-intelligence-driven triage, combining targeted edge defense with automated, agentic remediation. For AI infrastructure specifically, it calls for immediate containment strategies, sandboxing autonomous agentic workloads, and risk-based vulnerability management. It also recommends that software providers run AI-enhanced code review before release, on the reasoning that if pre-release AI review becomes standard, disclosure growth could slow."
      }
    },
    {
      "@type": "Question",
      "name": "How does Anomity help with this?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "The Endpoint Sensor inventories AI software across the fleet: 144 tracked AI tools, plus the MCP servers, CLIs, coding agents and local LLM runtimes that GTIG's largest categories describe, with versions. That is the asset list a GTIG-style triage process needs before it can prioritize anything, built from what is actually installed rather than what was procured. Findings route to SIEM, Slack, email or Jira, and Anomity complements existing vulnerability management rather than replacing it, by supplying the AI-specific inventory it has been missing."
      }
    }
  ]
}
```

```json
{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    {
      "@type": "ListItem",
      "position": 1,
      "name": "Home",
      "item": "https://anomity.ai/"
    },
    {
      "@type": "ListItem",
      "position": 2,
      "name": "Blog",
      "item": "https://anomity.ai/blog/"
    },
    {
      "@type": "ListItem",
      "position": 3,
      "name": "GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks",
      "item": "https://anomity.ai/blog/gtig-ai-software-vulnerabilities-agent-frameworks-2026/"
    }
  ]
}
```
