This same statement has been said countless times, replacing “AI” with whatever trend is big at the time. It has also been wrong in every case where that thing is said to replace QA.
The article makes a strong point that the slice that needs humans is dwindling quickly. The article seems to suggest our ability to write prose for other humans is our main differentiator.
just from first principles its really not. given its propensity to just make shit up and cheat, ai is really useful when there is some kind of formal or exhaustive checking as a wall for it to throw crap at. so if you rigorously define success and spend a bunch of tokens, you could easily save time and money. but if you ask the ai 'is this correct' and it says 'yes!', or even 'no!', then you've really learned nothing. this isn't a can you can kick arbitrarily far down the road.
I think writing tests is a great use of ai, but only if the tests themselves are throughly reviewed or are themselves validated by statements in a formal system.
Say what you like about the random word sequence generation machine, but it does actually seem to be usefully good at generating sequences of random words that correspond to problems in your software. And if you're inclined to write the code by hand, it's probably going to be easier to fix up your existing shit than rewrite it all. (And if you're going to use AI, then you're hardly going to listen to me.)
The purpose of this is to address a practice common in certain open source ideologies of banning all AI contributions - including vulnerability reports. Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.
> Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.
That's not only an uncharitable take, it's also wrong.
If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.
I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.
The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?
> The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?
Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).
See my reply to a sibling poster who also implies I did not read the article.
I think the disagreement here is whether this constitutes a "specific source" or not. I've seen some people produce things with incredibly quality and others produce useless slop all with the same AI tools and models, so defining that as one single source doesn't seem like a very good way of viewing things.
Lets assume your argument is valid, and further assume hat only the high-skill people submit AI-reports.
Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.
The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.
You could make the same argument for rejecting all reports from third-parties in a pre-AI world though. It just seems like you're picking an arbitrary property to extrapolate a trend from when there are plenty of other similarly arbitrary properties you could extrapolate similar trends from. If I noticed that most bug reports that come on Tuesdays are low quality, I don't think it would make sense to have any bugs that get reported on Tuesdays get auto-closed, regardless of the statistical trend.
Projects were already doing this by making you jump through hoops to submit an issue. It filters out the low effort chaff. I think asking for people (instead of their agents) to submit reports is on the same level.
If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:
> Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)
> If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:
I read the article very carefully, including the bit that you quoted. Here's what I read:
> Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.
And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.
I mean, he even said:
> I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:
Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.
(Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)
Yes. They will need to use AI to process the increased load of AI reports.
I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...
"use the tool to fix issues created by using the tool" is not a valid solution. The solution is to stop using bad tools.
> I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.
When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.
You are out of touch unfortunately. Your logic was correct last year, but things have changed radically and super fast. You need to re-evaluate, and this article by a core GNOME developer specifically doing security work should have made you do that re-evaluation!
The last 1000 times someone has said "nah bro AI was bad last year but this year it's good trust" have been wrong, I am also comfortable being informed by evidence.
They did this the last 999 times and always came to the same conclusion: that the AI promoters were talking nonsense. At what point should you stop listening to people who only talk nonsense so far, to avoid getting DoS attacked? Must the villagers look for the wolf every time the boy cries?
No. But when the wolf experts announce that there are a dangerous number of wolves - and you ignore it - that's a problem.
The people writing this article are experts. They cite other experts.
If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.
> The people writing this article are experts. They cite other experts.
"Economists have predicted 18 of the last 2 recessions".
I mean, c'mon! You have never read that?
Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.
In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".
You can do this by vibe coding extensions, and it appears many people are doing this. Ironically this probably makes GNOME the easiest DE to customize at this point.
They should be using an automated AI agent to validate vulnerability reports. Have it pull up the code base, confirm the bug exists, try to reproduce, then update the ticket. You might even run another AI agent pass to clean up the text before humans look at it. AI writing quality improves a lot with multiple passes.
"Vibe reviewing" is an underrated way of dealing with vibe coded PRs. There's going to be a lot of low-quality noise that's probably better to just close without a human looking at it, but it would be unfortunate to ignore the stuff that might be worthwhile because of that.
The era of software quality has always been there like in projects like postfix. Gnome could reduce features and increase testing and auditing.
But let that not waste the opportunity to promote AI.
AI is well-suited to the tedious work of testing and QA.
This same statement has been said countless times, replacing “AI” with whatever trend is big at the time. It has also been wrong in every case where that thing is said to replace QA.
"of testing and QA" ... for a certain percentage of "testing and QA". The rest still needs humans.
The article makes a strong point that the slice that needs humans is dwindling quickly. The article seems to suggest our ability to write prose for other humans is our main differentiator.
just from first principles its really not. given its propensity to just make shit up and cheat, ai is really useful when there is some kind of formal or exhaustive checking as a wall for it to throw crap at. so if you rigorously define success and spend a bunch of tokens, you could easily save time and money. but if you ask the ai 'is this correct' and it says 'yes!', or even 'no!', then you've really learned nothing. this isn't a can you can kick arbitrarily far down the road.
I think writing tests is a great use of ai, but only if the tests themselves are throughly reviewed or are themselves validated by statements in a formal system.
> given its propensity to just make shit up and cheat
That's more-or-less how I define QA work. The goal is not proving overall correctness, it's surfacing individual issues.
Everything coming from GNOME about software quality should be taken with a Strategic Petroleum Reserve of salt.
While other projects rewrite things in memory-safe ways, GNOME's response is to ask a chatbox if there are any memory vulnerabilities.
Say what you like about the random word sequence generation machine, but it does actually seem to be usefully good at generating sequences of random words that correspond to problems in your software. And if you're inclined to write the code by hand, it's probably going to be easier to fix up your existing shit than rewrite it all. (And if you're going to use AI, then you're hardly going to listen to me.)
I'm a KDE guy through and through, don't get me wrong, but statements like this deserve attribution, otherwise it's just FUD.
The purpose of this is to address a practice common in certain open source ideologies of banning all AI contributions - including vulnerability reports. Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.
> Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.
That's not only an uncharitable take, it's also wrong.
If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.
I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.
The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?
> The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?
Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).
See my reply to a sibling poster who also implies I did not read the article.
I think the disagreement here is whether this constitutes a "specific source" or not. I've seen some people produce things with incredibly quality and others produce useless slop all with the same AI tools and models, so defining that as one single source doesn't seem like a very good way of viewing things.
Lets assume your argument is valid, and further assume hat only the high-skill people submit AI-reports.
Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.
The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.
You could make the same argument for rejecting all reports from third-parties in a pre-AI world though. It just seems like you're picking an arbitrary property to extrapolate a trend from when there are plenty of other similarly arbitrary properties you could extrapolate similar trends from. If I noticed that most bug reports that come on Tuesdays are low quality, I don't think it would make sense to have any bugs that get reported on Tuesdays get auto-closed, regardless of the statistical trend.
Projects were already doing this by making you jump through hoops to submit an issue. It filters out the low effort chaff. I think asking for people (instead of their agents) to submit reports is on the same level.
If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:
> Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)
> If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:
I read the article very carefully, including the bit that you quoted. Here's what I read:
> Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.
And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.
I mean, he even said:
> I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:
Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.
(Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)
Yes. They will need to use AI to process the increased load of AI reports.
I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...
"use the tool to fix issues created by using the tool" is not a valid solution. The solution is to stop using bad tools.
> I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.
When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.
> When LLMs actually can reliably do their jobs
LLM agents do a fantastic job of finding exploitable bugs in code. MUCH better than humans.
They would also do a fantastic job isolating the duplicate reports as described above.
So what's your issue?
You are out of touch unfortunately. Your logic was correct last year, but things have changed radically and super fast. You need to re-evaluate, and this article by a core GNOME developer specifically doing security work should have made you do that re-evaluation!
The last 1000 times someone has said "nah bro AI was bad last year but this year it's good trust" have been wrong, I am also comfortable being informed by evidence.
> I am also comfortable being informed by evidence.
Then look. If you can't judge, then trust the experts. This article is written by experts.
They did this the last 999 times and always came to the same conclusion: that the AI promoters were talking nonsense. At what point should you stop listening to people who only talk nonsense so far, to avoid getting DoS attacked? Must the villagers look for the wolf every time the boy cries?
No. But when the wolf experts announce that there are a dangerous number of wolves - and you ignore it - that's a problem.
The people writing this article are experts. They cite other experts.
If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.
What if the last 999 times wolf experts announced there were a dangerous number of wolves, no wolves were found?
You're going to need to leave the hyperbole behind and use actual facts if you want to continue.
Cite data from the last few months.
But you get to use hyperbole to defend your side? No, I'm not interested in fighting Brandolini's Law.
> The people writing this article are experts. They cite other experts.
"Economists have predicted 18 of the last 2 recessions".
I mean, c'mon! You have never read that?
Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.
In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".
Your argument boils down to "I'm not trusting that bridge! Have you seen how many mistakes astrologers make!!"
I've found that software engineers are good at analyzing the public reports they receive for their own projects.
Let's not over generalize, shall we?
Hey does that mean they'll use "AI" to allow users to customize the desktop environment again?
You can do this by vibe coding extensions, and it appears many people are doing this. Ironically this probably makes GNOME the easiest DE to customize at this point.
> GNOME the easiest DE to customize at this point.
It's not even false, you do know that KDE exists
No - non-customizability is a design goal of the GNOME desktop, since they want the experience to be identical for everyone.
They should be using an automated AI agent to validate vulnerability reports. Have it pull up the code base, confirm the bug exists, try to reproduce, then update the ticket. You might even run another AI agent pass to clean up the text before humans look at it. AI writing quality improves a lot with multiple passes.
"Vibe reviewing" is an underrated way of dealing with vibe coded PRs. There's going to be a lot of low-quality noise that's probably better to just close without a human looking at it, but it would be unfortunate to ignore the stuff that might be worthwhile because of that.
Sounds really expensive as far as the work you're willing to let the public invoke on your project.
> AI writing quality improves a lot with multiple passes.
Or you end up with an AI version of Chinese whispers.