Written by Jeremy Souffir Founder, JTS Tech Services

The short version, and the direct answer if you read nothing else. On 30 September, Google's Threat Intelligence Group published an analysis of how AI has changed vulnerability discovery and exploitation, drawn from the twenty months between 1 January 2025 and 31 August 2026. Three findings in it matter to anyone running automation in a business. First, the raw volume has roughly doubled: 5,045 vulnerabilities disclosed in January 2026, 10,477 in July, 10,740 in August. Second, the flaws being found by AI research agents are a different and nastier shape — exactly 50% of them result in remote code execution, against 26% across the broader CVE ecosystem. Third, and this is the one worth stopping on: of 2,076 cumulative AI-related vulnerability disclosures, orchestration middleware accounts for 50% of all AI-related flaws, with a 347% surge in disclosures during 2026. Orchestration middleware is not an exotic category. It is the glue — the layer that joins a language model to your email, your spreadsheet, your database and your store. If anyone in your business has built an automation in a visual workflow builder in the last year, that is the software this report is about. And the practical change is not the count. It is the clock: on one flaw Google tracked, a threat cluster was exploiting it four days after public disclosure, with five more clusters inside seven days.
Whose figures these are, and what this report does and does not claim
- The source is Google's own threat intelligence team, writing about data it collects. That is about as good as provenance gets in this field, and it is linked at the bottom. It is also not a neutral party: Google sells vulnerability management and AI security products, and the report closes by recommending them. The figures are measurements, the recommendations are marketing, and it is worth reading the two differently.
- The window is 1 January 2025 to 31 August 2026. Anything after that is not in these numbers, including the whole of September. So treat the monthly disclosure counts as a trend that was still climbing when the data stopped, not as a current reading.
- The disclosure counts are not a count of attacks. More than ten thousand vulnerabilities disclosed in a month does not mean ten thousand things are being exploited. The report is specific about that distinction and gives the exploited figure separately: 127 vulnerabilities were exploited across the whole of 2025, against 141 distinct vulnerabilities disclosed and exploited in the first eight months of 2026 alone. That is the number that should worry you, and it is three orders of magnitude smaller than the disclosure count.
- The orchestration finding is about disclosures, not about your installation. Half of AI-related flaws being in the glue layer tells you where researchers are finding problems. It does not tell you that your particular setup is vulnerable, and nothing in this post should be read as saying it is. It tells you which part of your stack has suddenly become the part worth checking first.
- What is not established is how much of this reaches small and mid-sized businesses. The report tracks disclosures and threat clusters, not victim profiles by company size. Anyone telling you this week what share of Canadian mid-market firms are exposed is guessing. We are not going to put a number on it either.

What is the orchestration layer, and do you actually have one?
This is the question that decides whether the rest of the report applies to you, and most businesses answer it wrong, because the honest answer is usually yes and nobody filed a ticket about it. Orchestration middleware is the software sitting between a language model and everything else. It holds the connections, the credentials, the sequence of steps, and the decision about which tool gets called with what. The report names the ones it tracked: Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, Letta, Pydantic-AI and the Model Context Protocol. Some of those are developer libraries that arrive inside a build. Several — Flowise, Langflow, Dify — are visual, drag-and-drop workflow builders, deliberately designed so that somebody who is not a software engineer can wire an automation together on a Thursday afternoon and have it running by Friday. That is their whole value proposition, it is a genuinely good one, and it is also exactly why this finding matters: the software with the highest concentration of new vulnerabilities in the AI stack is the software most likely to have been installed by somebody who was not thinking about patching it.
The mechanism is worth understanding precisely, because it is unusually specific and it explains why this category and not another. Visual workflow builders almost always include a node that runs code — a Python step, a custom function, a shell call — because without one the tool cannot do the long tail of things people need. That node is the point. The report describes attackers reaching those code-execution nodes either through prompt injection or through a crafted workflow file, and its phrasing for the result is the clearest sentence in the whole document: turning natural language prompts into unauthenticated remote code execution. Not remote code execution that needs a password. Unauthenticated. The named examples are concrete enough to check against your own installation.
- CVE-2025-3248 in Langflow: unauthenticated Python code injection through an exec() call in a code-validation endpoint, giving immediate remote code execution. The report lists this one as having been exploited in the wild.
- CVE-2026-5027 in Langflow: a path-traversal file write in the file upload handler, letting a remote attacker drop files of their choosing onto the host — the report's examples are cron jobs and SSH keys, which is to say persistence rather than a one-off.
- CVE-2026-42271 in BerriAI LiteLLM: command injection in a Model Context Protocol server's connection-test endpoint, resulting in host takeover and theft of the API credentials held there. If one box holds the keys to every model and tool your automations touch, that is the box.
- The pattern across all three: the vulnerable thing is not the model and not the business system. It is the component in the middle that holds credentials for both ends and accepts input from one of them.
The detail that changes the arithmetic
Pay attention to the risk-profile finding, because it is the part that quietly invalidates how most organisations triage. Across the broader CVE ecosystem, 26% of vulnerabilities lead to remote code execution. Among vulnerabilities discovered by AI research agents, exactly 50% do. The report also notes the distribution shifting upward in severity: 58% of AI-discovered flaws qualify as medium threat risk, more than double the baseline, while the low-risk share falls to 39%. If your process for handling an advisory is to look at the volume, assume most of it is noise, and wait for something marked critical, that process was calibrated on a population where three quarters of findings could not run code. The population changed. Nothing about this says every advisory is now urgent — the report does not claim that and neither do we. It says the base rate you were implicitly using is no longer the right one, and a triage rule tuned to the old base rate will now let through roughly twice as many code-execution flaws as it used to.
Why four days matters more than any of the volume numbers
The single most useful thing in the report is one case, not a statistic. CVE-2026-1731 is an unauthenticated operating-system command injection flaw in BeyondTrust's Privileged Remote Access and Remote Support products. It was found autonomously by a third-party research agent, built by a company called Hacktron AI — not by a human researcher who then wrote it up. Within four days of public disclosure, Google observed a threat cluster exploiting it. Within seven days, five more clusters were on it. The follow-on activity was the ordinary grim list: privilege escalation, data exfiltration, and secondary payloads dropped onto the compromised hosts. Read that sequence again with your own patching cadence in mind, because that is the comparison that matters. Four days is shorter than a lot of monthly maintenance windows. It is shorter than a change-approval cycle at most businesses that have one. It is shorter than the average summer holiday.
And the reason it compressed is structural rather than a one-off. When a human researcher finds a flaw, there is a natural delay built into the whole chain — the finding takes weeks, the write-up takes days, and the attacker who reads the advisory has to do their own work to weaponise it. When research agents find flaws at machine pace, and attackers also have agents reading advisories and building against them, both ends of that delay shrink at once. The report's framing of this is that defenders now face compressed adversary timelines, and its recommendation is to modernise triage rather than to patch faster, which is the right distinction. Patching faster is a capacity problem you probably cannot solve. Knowing which three things to patch first is a knowledge problem you can.
- A monthly patch cycle is now, on this evidence, slower than the exploitation window on a serious flaw in software you run. That does not mean move to daily patching; it means the cycle can no longer be the only mechanism.
- You cannot triage an inventory you do not have. The thing that makes four days survivable is knowing, on day one, whether you run the affected component at all — and most businesses cannot answer that about their automation tooling in an afternoon.
- The components most likely to be affected are the least likely to be in your asset register, because they were installed to solve a workflow problem rather than through a procurement process.
- Internet exposure is the multiplier. An orchestration server reachable from the public internet with an unauthenticated code-execution flaw is a different category of problem from the same software bound to localhost behind a VPN. Both appear in the same advisory, and the advisory will not tell you which one you are.

Isn't this the same point we made on 30 August?
Fair challenge, and worth answering directly rather than hoping nobody notices. At the end of August we wrote that AI agents do not need zero-days — that an agent had read its own host's kernel version, fetched a published exploit for an unpatched flaw, and become root, and that the lesson was to close the known holes rather than to worry about the exotic ones. That argument still stands and this report reinforces it. But it took the set of known holes as a given, and that is precisely what has changed. This report is about the pipeline producing that set: it got roughly twice as large per month over 2026, it tilted sharply toward flaws that permit code execution, and it concentrated half its AI-related output in one layer of software that barely existed in most businesses two years ago. In August the instruction was patch the known flaws. The addition now is that the known flaws arrive faster, hit harder, and land disproportionately in the part of your stack you are least likely to have written down. Same direction, different problem — one is about diligence, this one is about inventory and timing.
The two conclusions that both get this wrong
Rip out the workflow builder and go back to doing it by hand
The reflex, and badly wrong on the arithmetic. The orchestration layer is where the flaws are being found because that is where the research attention went this year, and research attention goes where adoption goes. A newly popular software category with a 347% jump in disclosures is a category being examined properly for the first time, which is how every widely used thing has looked at some point in its early life — web frameworks, container runtimes, CMS plugins, all of them. The flaws being found are, on the whole, being found by people who report them. Tearing out an automation that genuinely saves your team eight hours a week, to avoid a class of risk you have not actually assessed in your own environment, trades a measured benefit for an unmeasured fear. Worse, it does not remove the risk — it moves the work back to humans doing it in email and spreadsheets, which has its own well-documented failure modes and no CVE list at all. The answer to software with vulnerabilities is patching and reducing exposure, which is the answer it has always been.
We're small, nobody is scanning for our Langflow instance
The opposite error, and the more dangerous one, because it feels like proportionate judgement rather than optimism. Mass exploitation of an unauthenticated remote-code-execution flaw does not involve anybody choosing you. It involves scanning the entire address space for the fingerprint of the affected software and firing the same request at everything that answers, which costs the attacker approximately nothing per host and is why the five extra threat clusters turned up inside a week. Your size is not an input to that process. Nor is your industry, your revenue, or whether anyone has heard of you. The only variables that matter are whether the software is reachable and whether it has been patched — and the post-exploitation payloads in the report include cryptominers, which exist precisely because an attacker with no interest whatsoever in your business can still make money from your server. Being too small to target and being too small to scan are completely different claims, and only the first one is true.
What is worth doing this week?
- Write down what orchestration software you actually run, by name and version. Flowise, Langflow, Dify, LangChain, LlamaIndex, CrewAI, AutoGen, Semantic Kernel, LiteLLM, any MCP server, any n8n or Zapier-style builder that executes custom code. Ask your developers, ask whoever built the automations, and check what is running on your servers rather than what you believe is running. This single list is most of the value in this entire post.
- For each one, establish whether it is reachable from the public internet. That one fact separates a routine patching task from something to deal with today. A management interface on a cloud instance with a public IP and no gateway in front of it is the specific configuration the mass scanning finds.
- Find the code-execution nodes and decide whether you need them. Most visual builders let you disable or restrict the step that runs arbitrary code. If none of your workflows use it, turning it off removes the exact path the report describes, and costs you nothing.
- Check where the credentials live. The LiteLLM case is a reminder that the orchestration box frequently holds API keys for every model and system it touches. Find out which keys are on that host, whether they are scoped or full-access, and whether you could rotate them this afternoon if you had to.
- Subscribe to the security advisories for the specific components on your list. Not a general feed — the actual project advisories, going to a named person's inbox. On a four-day window, hearing about a flaw at all is most of the battle, and nobody hears about it from a dashboard they do not read.
- Decide now who patches an orchestration component out of cycle, and who may authorise taking it offline. Make that decision while nothing is wrong. The reason four days is dangerous is almost never that the patch was hard; it is that the first two days went on working out whose call it was.
- Do not try to produce a risk score for this. Somebody will ask how exposed the business is, as a number. You cannot know that before the inventory exists, and a confident figure invented this week will be quoted back at you for a year. Saying the list is being built and the exposure question comes after it is a better answer.
The genuinely encouraging part
Almost everything this report asks of you is cheap, and none of it needs a budget approval. The expensive version of AI security is a platform, a vendor and a programme; the version that would actually have saved the victims in this report is a list of software names, a check on which of them answer from the internet, a decision about one toggle, and a named person who reads four advisory feeds. That is an afternoon of work and it holds its value for a year. There is a second piece of good news hiding in the numbers, too. The reason the disclosure count doubled is that AI is now very good at finding vulnerabilities — and that capability is not exclusively available to attackers. The same report notes the defensive side of it, with automated code review catching flaws before software ships. The flaws being counted here are overwhelmingly flaws that got found and published, by researchers and research agents working in the open, which is the system functioning rather than failing. A year ago, a path-traversal bug in a workflow builder's upload handler might have sat there unnoticed. Now it gets found, written up and patched — and the only thing standing between that patch and your server is knowing that the software is yours. That part is entirely within your control, and it is the one thing no vendor can do for you.
Where we fit
The reason this particular gap opens is that automation tooling enters a business through the side door. Nobody procures a workflow builder. Somebody has a repetitive job, finds a tool that solves it in an afternoon, and the automation becomes load-bearing long before anyone decides who owns the server it runs on. By the time an advisory lands, the person who built it has moved on or moved roles, the credentials on the box are undocumented, and the four-day window is spent establishing basic facts rather than applying a patch. We do two things about that, and they are deliberately unglamorous. The first is the inventory: every integration, scheduled job, orchestration component and standing credential that touches your systems, written down with a named human owner, a version, and an honest answer about what is reachable from outside — which is the artefact that turns a frightening advisory into a ten-minute check. The second is what happens when the check comes back badly. When a system is actually under attack or behaving strangely, our AI agents run the forensics across server logs, edge analytics and configuration in parallel while a senior engineer directs the investigation and makes every remediation call, fixes get verified with live tests, and everything lands in a change ledger with rollback paths. The reason to have that relationship in place beforehand rather than afterwards is simply timing: on a four-day exploitation window, the time you spend explaining your architecture to somebody new is time you do not have. Our bot-swarm case study is the honest version of what that looks like in practice, and it is linked below rather than summarised here.
Sources
- Google Cloud / Google Threat Intelligence Group — Vulnerability Discovery and Exploitation Trends in the AI Era (the primary source, published 30 September 2026, covering 1 January 2025 to 31 August 2026: the monthly disclosure counts of 5,045 in January 2026, 10,477 in July and 10,740 in August; 127 vulnerabilities exploited across all of 2025 against 141 disclosed and exploited in January–August 2026; 2,076 cumulative AI-related CVE disclosures with over 1,500 in 2026 alone; orchestration middleware at 50% of all AI-related flaws with a +347% surge; the 50% remote-code-execution rate among AI-discovered vulnerabilities against 26% across the broader ecosystem; the 58% medium-threat-risk share; the named frameworks; the prompt-injection and crafted-workflow-JSON path into code-execution nodes; CVE-2026-1731 in BeyondTrust found autonomously by Hacktron AI with one threat cluster exploiting it within four days of disclosure and five more within seven; the SNOWLIGHT, SPARKRAT and cryptominer payloads; and CVE-2026-42271, CVE-2026-5027 and CVE-2025-3248)
- Help Net Security — The vulnerabilities AI finds are the ones attackers want (independent reporting on the same report, 1 October 2026, for the remote-code-execution profile and the significance of the orchestration-framework concentration)
- SecurityWeek — Google: AI Is Changing the Pace and Profile of Vulnerability Discovery (second independent write-up, for the pace-and-profile framing and the exploitation timeline)
- Infosecurity Magazine — AI-Found Vulnerabilities More Likely to Enable RCE, Google Says (third independent write-up, for the severity-distribution figures)
- SiliconANGLE — Google finds vulnerability disclosures doubled as AI changes which flaws get discovered (same-day coverage, 30 September 2026, for the doubling of disclosure volume)
- NIST National Vulnerability Database — CVE-2025-3248 (the Langflow code-injection entry, if you want to check the affected versions against your own installation rather than take anyone's summary for it)
- JTS Tech Services — Agents don't need zero-days (the August post this one extends: an agent using a published exploit against an unpatched host, and why known flaws are the ones that matter)
- JTS Tech Services — A poisoned package was live for forty minutes in March. In August, the keys it stole still work. (on credentials left unrotated after an incident in the AI tooling layer, which is the same box this report puts at the centre of things)
- JTS Tech Services — AI-accelerated incident response: containing a bot swarm (what the forensics described above actually looked like on a real incident, including what was fixed and over how long)


