Research indicates that there is a potential issue lurking beneath the surface of modern software—significantly more code is now being written by AI and a large percentage is not being discovered as security vulnerabilities. Organizations are piling up hidden risks that could emerge as breaches as development teams increasingly rely on AI coding assistants, security researchers warn.
What the Research Shows
The concern is not hypothetical. The AI-generated code is four times more likely to contain security flaws than hand-written code, according to Gartner. The answer is due to the way the tools function. The code that an AI assistant predicts the most likely restatement of is full of insecure, outdated and flawed examples, while this enormous amount of publicly accessible code it learned from is. The model does so at scale, generating code that seems legitimate and performs as expected, but has the same flaws as the code used for training.
Worse, researchers have found a behavioral risk factor that they say is particularly worrisome. AI assistants have been found to increase the likelihood of developers writing insecure code and feeling more confident about the security of their code. The suggestion originates from a machine, a fact that implies authority and thus those security-critical lines are slotted by thorough review that a developer would have more carefully considered if the line had been written by a colleague.
Why Volume Turns a Small Rate Into a Large Problem
When you have a lot of code, a small vulnerability rate is a big problem. AI assistants can produce thousands of lines of code in minutes that a developer might take many hours to write and adoption is on the rise among enterprise engineering teams. The further the total volume of code goes into production, the more absolute flaws are introduced, quicker than traditional review processes were ever intended to be able to catch them.
This is where the strain shows. Conventional application security processes were built for a slower era, when code changed on a predictable schedule and a periodic assessment could reasonably capture the state of a system. The thing is, with continuous, AI-powered development, that’s no longer the case. Quarterly reviews mean nothing about the thousands of lines that might have been shipped in between, and this becomes an ever-bigger gap between tested and running code.
The Limits of Existing Tools
Many organisations are filling that gap with static scanners, but researchers point to problems with static scanners, too. They don’t run code to check if there is a real vulnerability, since they are analyzing code rather than running it. This translates to a false positive rate of up to 60%, drowning security teams in alerts that often have limited use. With that noise, teams end up spending precious hours looking for smoke and mirrors or tuning out of the scanner completely, both of which are ways that real vulnerabilities make it to production.
Adding to the issue is the lack of consistency when it comes to the AI-powered reviews. The large language models produce results in a probabilistic rather than deterministic fashion, leading to different results on two different scans of the same code, with vulnerabilities identified on one scan and not the other. Such variability is a problem for security teams looking for consistent and reliable outcomes, because it threatens the trust they place in the assessment.
A Different Approach to Securing AI Code
The answer that is emerging across the industry is: Test the running application, not the static source; validate all results before they get to a developer. Platforms offering AI-powered AppSec are built around this principle. Bright Security’s STAR platform that identifies, fixes, and verifies real vulnerabilities quickly and early in the software development lifecycle—regardless of whether they are hand-crafted or generated by AI.
Instead of providing teams with a list of theoretical issues, it validates which ones can be exploited by exercising the live application, and then provides verified fixes, automating up to 98% of remediation. It validates against the running system, ensuring false positives of less than 3%, and speeds up vulnerability resolution by up to 10x. The pairing enables AI-generated code to be tested and secured at the same rate it’s created by developers within the pipelines and repositories where they already operate as opposed to in a distinct review that happens too late to matter.
What It Means for Organizations Shipping AI Code
The researchers themselves are not saying that people should avoid using code that has been created with AI. The results in terms of productivity are tangible and adoption is expected to continue. The cautionary note is that sending that code without some testing for speed and volume involving security measures is a risk and one that will slowly creep up until it’s an incident.
Organisations that are already producing a significant percentage of their software using AI can take nothing more than what you already know. Security must be running all the time, verifying the results it receives and remediating quickly enough that it can outpaced by development. The vulnerabilities that the researchers are talking about are only hidden until someone looks, and it’s the organizations that are testing every build automatically who are closing that gap, not the ones that assume that the code written by the AI is safe because that’s what it looks like.
Media ContactCompany Name: Bright SecurityContact Person: Eyal DrorEmail: Send EmailPhone: 546526651Country: United StatesWebsite: https://brightsec.com/