I see symptoms of this all the time. For example, it's a weekly annoyance for folks to pop into /r/strava to showcase their vibe-coded app that uses the Strava API to do $THING. Then someone invariably points out that an existing app (or even Strava itself) already does $THING, and often it's free. I don't mean to be negative, I think it's great that people are building useful niche software and I don't blame them for wanting to share it. A significant part of the problem is that it's much harder nowadays to find "prior art" because keyword/boolean web searches have been FUBAR.
Hopefully now that AI provides quick answers, web search can go back to providing and respecting boolean and more advanced techniques. If web search is no longer the go to tool for most users, then search can do better what it can do uniquely?
I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
I called it 2001/2002 or whenever they appeared when I tried to explain why personalized search results are the beginning of the end of a shared reality and therefore the ability to reason and act in public, and with others. I bet some still consider it hyperbole. It's just taking in trends and seeing where the glacier moves to, how the cookie will crumble so to speak.
It is an astute point. I agree that we have largely lost a lot of our shared reality. But as with any "beginning" of a long, diffuse process, people might quibble about the specifics.
I offer a few other moments you might gesture toward as the beginning of the end of shared reality.
First, 1987, with the elimination of the Fairness Doctrine and the subsequent boom in partisan talk radio shows.
Second, 1989, when cable TV became a mature technology reaching half of U.S. households.
And finally, maybe a lesser item, the introduction of the DVR circa 2000, when we stopped watching scheduled TV together.
I wonder if more historically informed people would find the fracturing going further back as other advancements in publishing technology allowed more people to share more views.
That was a temporary blip. Specific technologies facilitated an age of mass, top-down communications in the mid 20th century, just by virtue of how those technologies worked. It enabled a small group of people in charge of the radio/TV stations to have a monopoly on communications and media. The technology both before (printing presses with myriad newspapers) and after (the Internet) were by their nature decentralized.
Not totally convinced about the printing press example (expensive, technical). The biggest presses, with the longest reach (daily newspapers) were often owned by the richest people.
What's unique now is decentralized + extreme potential reach. Or it was. Now we have to consider bias in the algorithm that sticks stuff in front of us and is, I suspect, the modern equivalent of those daily newspaper presses.
But corporations were an act of Congress in 1776. Early America was decidedly anti-corporate (the Boston Tea Party was about a tax break to a corporation). Early Americans (and early thinkers who influenced Americans) were grossly distrustful of corporations and pools of money.
> The directors with of such [joint-stock] companies, however, being the managers rather of other people's money than of their own, it cannot well be expected, that they should watch over it with the same anxious vigilance with which the partners in a private copartnery frequently watch over their own
The early pre-United States had pamphlets and newsletters. Like, a lot of them. It was a big influence on the early post office. You could get rich running a printing press but you did it by printing for the masses.
The yellow journalism era, on the other hand, was a little closer to the rich owning the press.
This discounting how many of the 100s of worker owned newspapers/magazines that were in circulation. Not everything was controlled by a select few. How did you think labor movements in America during the 1800s collectively organized across the country sharing literal war stories?
I'm not discounting it by any means. And I know about the pamphlet era that was enabled by the invention of the printing press, and the profound effect it had on Europe. But what's the typical reach of one of those worker-owned newspapers, compared to the reach of a mass circulation newspaper?
I think you missed my point: printing at scale needs a support system and a distribution system, and that requires capital. The "age of mass, top-down communications" didn't start "in the mid 20th century", and the printing press should not be used as a counterexample to support rayiner's argument. William Randolph Hearst was born in 1863.
"Oh those don't count as newspapers. Newspapers are professional. Those are just zines."
It's why I hate people looking down on fanfiction, too. Aliens was Alien fanfiction. The New Testament is OT fanfiction. Hell, nolan's The Odyssey is fanfiction.
The establishment uses terms to dismiss outsider art as less-than, and this isn't much different. (They of course get to decide what the "inside" is)
It takes some technologies more time than others to succumb to the corrupting pressures of money, but they all fall eventually. We must keep outrunning it, building ever more resistant methods of communication.
I'm hopeful that the next one will last even longer since we'll have to build it explicitly for resisting corruption. Neither the printing press nor radio nor webs 1 or 2 had fighting back as part of their DNA. 3 was a bit of a flop, but there's a lot of design space still out there to explore.
The top-down centralization of mid-20th century technologies wasn't caused by "money," it was caused by their technological nature. They inherently relied on limited radio spectrum that was doled out to specific companies by the government.
A fully traceable public ledger is kind of a wet dream for an intelligence agency/evil corp. Funny how much faith we put into that, and look at it now.
What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
And even when TV was limited to a small few national OTA (over the air) channels, different people would tune in at different times to watch their different preference of shows. So just because content was fewer, it didn’t necessarily equate to different people consuming the same content.
The real crux of the modern problem isn’t preference drive choices. It’s independent publishers being driven out by corporate greed. Eg AI traffic making it unviable for independent blogs. But then one could argue that the current ecosystem, where the barrier for publishing being so low, is an anomaly because historically that was always prohibitively expensive. To go back to the TV example: you couldn’t commission a show without deep pockets and a lot of TV exec contacts.
I’d love our golden age of information to be persistent. But, and as yourself and others have alluded to, that’s not guaranteed unless we fight to keep it that way.
>What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
Exactly. The shared reality is the experience. You both were given the same opportunity to take in the same information. The different opinion is what comes out of the other side of the meat suit processing the experience.
Algo driven results changes the experience by changing what the user perceives. So there might never be a shared experience.
I think of it more as a basis for shared culture - I could chat about Seinfeld / Friends, and odds were high that the other person would know what I was talking about.
I imagine in prior ages this was perhaps more about folks reading a core collection of books. I expect there was less shared culture than in ^
Now it seems like we're in an era with (potentially) less of a shared cultural bases than either of those times. I've certainly met people where there's really no overlap - their core beliefs, values, etc - fundamentally different foundations. If that becomes more widespread, it's hard to know what that will do to our societies.
The thing is, people I’m friends and thus likeminded, we watch the same shows and still have those kinds of conversations. And people I wasn’t likeminded with never used to watch the same shows, even pre-streaming era, so I’d never had those conversations.
Take your examples, Friends and Frasier were aired at roughly the same time in the UK (obviously on competing stations). Most people watched Friends, but that show wasn’t really my thing. So I watched Frasier. As did my closest friends. And we’d chat about that. But I couldn’t join in with work colleagues with Friends chats.
It’s a bit like sports matches. People at work will talk about football/soccer but I’d be more interested in Snooker. So I’d chat snooker with friends and duck out of the football work chat.
Independent blogs are perfectly viable. It is basically free to publish. What's not viable and never was is to be a professional blogger. At best, some people could try to be professional shills with blogs as bait (it is good if these go away; they were noise), and an extremely tiny minority might actually make a living with people paying to support their blog directly.
I think being a professional blogger is on the rise, the driver for the viability of being a pro blogger has risen. Substack/medium etc. streamline the getting paid part. The fracturing social narrative, (rise in censorship, cancel culture, tribalism) actually drive people to pay to gain access to periodtical writing from specific people seen as authorative on one or more subjects. These subjects that mainstream media are too shy to surface, too dry to feature, or have a tendency to twist to a narrative that supports the sponsors etc.
> What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
The shared reality in that example includes only facts: OP and mum went to a movie. The movie was called X. The actors were Y and Z.
The original point of this subthread was that we can't even agree on a shared set of facts anymore! The political Right has an entirely different set of facts than the political Left, and when our reality is derived from those facts, we cannot share the same reality anymore.
Then, we're just back to OP's conundrum: "Everything is opinion." I'll say grass is green and you'll say grass is red, and there can be no fact because they differ.
Not really no. Some things can be proven while some cannot.
The colour of grass is something that’s scientifically measurable.
My point is that some broadcasters are very strict about sticking to the facts. Some will misrepresent those facts but not technically lie (eg discuss statistics that favour their partisan view but don’t share the full context behind those stats).
And then there’s broadcasters like Fox News that promote flat out lies. It’s often not even an interpretation of the truth. It’s shared as an opinion but the evidence clearly proves it’s just lies presented as facts.
And I really wish there were criminal charges for broadcasting blatant lies. Not just in America, but most of the developed world where populist-propaganda is used to brainwash gullible voters.
Broadcasters absolutely have a responsibility to ensure they communicate truthful statements in exactly the same way that any other profession has a responsibility to ensure their output is correct.
Everything is perceived, and our perception gets filtered based on our prior experience. Just because we're not directly perceiving objective reality, that does not mean _everything_ is opinion.
JFK was shot, that's not just my opinion.
Unless you choose to believe some kind of post-modern crazyness, in which case - have at it.
You don’t actually know JFK was shot from your own personal experience. You just choose to believe what credible source claim. And even if you were there, who’s to say that your fallible memory is incorrect?
To be clear, I don’t think this line of thinking helps much. But it’s an interesting philosophical dilemma.
"Who’s to say that your fallible memory is incorrect" - you can very easily test if someone has a bullet through their head. Well, assuming someone didn't go and steal their brain.
I'm saying that not everything is opinion, that there does exist a category called 'fact'. And you're talking about something else - a chain of trust with respect to those facts.
This is starting to go down the crazyness line I had mentioned. I get it's fun to philosophize about such things, but it's not how we live.
> you can very easily test if someone has a bullet through their head. Well, assuming someone didn't go and steal their brain.
Do you have access to their head? How do you know that someone didn’t swap that head for someone else who had been shot?
Like I said, this is purely a philosophic argument. I’m not actually trying to suggest that there isn’t such thing as “facts”.
> And you're talking about something else - a chain of trust with respect to those facts.
No, I’m making a psychological point that our perceptions are what we use to construct our reality, and that our perceptions are malleable. If they weren’t, then drugs would have no effect, psychosis wouldn’t be a thing, and people’s opinions wouldn’t differ about even the most straightforward things like god, globe vs flat earth, and so on and so forth.
> This is starting to go down the crazyness line I had mentioned. I get it's fun to philosophize about such things, but it's not how we live.
I agree. As I said in the previous comment:
“To be clear, I don’t think this line of thinking helps much. But it’s an interesting philosophical dilemma.”
There is an pink elephant behind you. In practice, I believe that at least after the dust have settled, events of the type of "Count Ferdinand was shoot dead infront of 100 witnesses", like maybe 1/50 are fake? Dunno what share.
The faking part would be in who is blamed and what country to bomb etc.
Not necessarily. A shared reality is one where people are discussing the same thing (within a reasonable margin) even if filtered by their individual perceptions, and drawing different conclusions.
I don't know the word for the opposite of shared reality (maybe private individualized realities), but it results in a scenario where you have nothing in common with other people other than "I saw some thing" and therefore it becomes very difficult to discuss anything with them.
I think different cultures can share the same reality, though filtered through different lenses.
If two people from different cultures watch the same movie, one finds it ok and the other offensive and disrespectful, that's a shared reality filtered through different cultural backgrounds. Even if either version of the movie is dubbed, censored, or altered in any way, there's still enough common or shared material that a discussion can be held.
If two people had completely different experiences and watched completely different, tailor-made movies, then there's very little for them to discuss or even disagree with. "I haven't watched your movie", "I haven't watched yours either", "Mine was better", "I guess... I wouldn't know, I thought mine was cool", etc. No shared reality.
Opinion and preference are the same thing. And thus personalised search results, choice of shows in the streaming era, and so on, are just a result of that personal bias.
And institutional control of information didnt really start again until the Statute of Anne struck up a deal between printers and the state giving printers copyright over written works and the state the right to cencor and ban works the disapproved of nearly 250 years latter.
Radio gave us Demagogues like Father Coughlin in the US and Hitler in Germany, and we still havent dealt with the problems radio introduced when television came along and now we have the internet with social media, podcast, and algorithmic echochambers, and algorithm induced radicalization. Radio television internet any one is just as much a shock to society as print was and we still havent figured out how to adapt yet to any fully.
> The industry boomed in the 1980s as more and more customers bought VCRs. By 1982, 10% of households in the United Kingdom owned a VCR. The figure reached 30% in 1985 and by the end of the decade well over half of British homes owned a VCR.
My impression of the VCR is not too many people managed to use it for DVR-style rescheduling live TV. Although it did have that feature. We mainly used it to watch movies that we had already been in the theater.
But, for sure, another fracturing. We weren't all talking about the nine movies in the theaters, but whatever we had rented and watched on tape instead.
My recollection was that the VCR's main selling point was the ability to record TV and play it back later, and I can say anecdotally, that that was the primary way my family used ours in the early 80s. Only later did the idea of Blockbuster and "using VHS merely as content distribution" take hold.
My family used it to capture programming we couldn't normally be around for, especially late night BBC content on PBS. Pretty much all we used it for was DVR style recording.
> But as with any "beginning" of a long, diffuse process, people might quibble about the specifics.
Yeah, absolutely. But for the web, at least for my daily use of it, this was the first big "shift". Learning how something was spelled by someone googling it and quoting the number of results to the others.. good times. And you can no longer find out what the page with highest page rank says about "pizza" regardless of where YOU happen to be.
Yes, there's certainly less of a monoculture now that people have greater choice in what media they consume. Personally I have trouble seeing that as a bad thing. Greater diversity of taste, culture, and opinion makes for a richer, more interesting society in my opinion. Though I recognize it does have downsides.
I wonder if there's a way to keep the positive effects of such diversity but mitigate the drawbacks. Certainly echo chambers are undesirable.
There are other people in those bubbles, it's not an isolated void. HN is one such bubble. Yes it's a smaller bubble, there's not nearly as many people here as there were watching broadcast television back in the 1980s, but it would be wrong to say there is no culture here. Same goes for the millions of other small communities out on the web and social media. But as I said, it's certainly true there are downsides associated with that diversity.
> I wonder if there's a way to keep the positive effects of such diversity but mitigate the drawbacks
The answer society mostly settled on is representative democracy.
You appoint qualified candidates via consensus, and then give them the microphone and some measure of control as long as the consensus stays strong. This is sort of what we have now with influencer culture.
The flaw has always been human nature. It turns out that half of all people are of below average intelligence and they consequently make bad decisions. Especially if the smarter half use them as pawns to assume control, which is what US politics has become.
Once the Morlocks and the Eloi have settled into their roles the entire thing falls apart and democracy dies screaming.
> It turns out that half of all people are of below average intelligence and they consequently make bad decisions.
Worse than that; it turns out that even people of above average intelligence consistently make bad decisions! AND they still attempt to control others!
Good point! Interested to hear about the RoW. Where are you from? Was there pre-custom search results fracturing or did you have a (relatively) uniform national media before that?
I never thought of the DVR as a particularly significant development. I thought of it as just a replacement for the VCR. That's what it was for me. But perhaps in some areas DVRs were used to a greater extent than VCRs had been?
Yeah, I readily concede it's lesser. For me it was the moment that we stopped watching a particular thing and discussing it the next day. But maybe that's more personal biography than sociology.
I've always thought that if online advertising as a business model had been entirely banned from the internet in the 90s, the internet would've looked like a much better place today.
And micropayments would've taken off in a big way.
The only reason why micropayments doesn't really exist today isn't due to technical reasons, it's because advertising is too dominant, and because consumers got trained on expecting everything on the internet to be "free", not realizing they're the product.
There are many indirect ways to get information though. Better would be to ban "discrimination" of customers/humans (no loyalty programs for long term customers/whales or people who know people), and ban of usage of all personal information for decision making except short pre approved lists. Most industries would do fine with "payment received", some maybe could make legitimate cases for also "delivery address" or perhaps "age". Would probably also be good for efficient competition on merit, which should appeal to the groups with enough wealth to use it to accumulate wealth.
Humans are only now figuring out pointy shoes are bad for people's feet. This species is wrong about almost everything when it comes to its own well being and like pointy shoes - it's a completely unnecessary self-own that barely benefitted anyone.
I think advertising runs a little deeper into our identity and history than you are thinking. If you have a sign in your window describing your wares, that's advertising. If you're standing on a street corner inviting customers to enter, that's advertising. If you're doing nearly anything to make the world know that your business exists, you're advertising. I'd rather not live in a world where "ad cops" see fit to intrude and opine on such activities. Attacking PII brokers, on the other hand, would do far less collateral damage.
Personalization came along in MMOs around the same time, and it was hard to communicate the difference to people between a persistent MMO zone as a "place" where people and things exist, even if those things are inappropriate for your character or hostile to the point that you can't survive there or too crowded for the intended flow... versus an instanced MMO zone, a little pocket universe for you specifically that has no strangers in it, which is more of an "experience" or "cut-scene" for you to imbibe prepared content.
A heavily instanced MMO may as well be a single-player or co-op game with a shared chatroom as a lobby, there is absolutely none of the emergent play, none of the characterization of players & groups of players, nothing original. Instead, give me the same ground to walk on as everybody else, even if our tread wears all the grass off of it; Half of my life in these games was getting bored waiting for a mob to spawn and fucking around making friends waiting in line, or ganging up on the guy who wanted to cut in line, or waging war on the clan that currently controls the needed territory, or grouping together to fight mobs higher than I should be able to hit (which the instance will just Adjust Downwards). Instancing and other auto-personalization algorithms cut the annoyances that are central to human socialization, achievement loops, and wayfinding.
Similar impact to non-massive multiplayer online games. Skill based matchmaking destroyed the third place that once existed on public game servers.
I still don't understand the appeal of an online game where you're algorithmically guaranteed a ~50% win rate after a calibration period. Seems to be selling well enough though.
I think the lure with matchmaking was that when it was introduced, it made you really good compared to before by easier access to better practice.
In Data 2 the skill level shifted upwards enourmously for pro players and I feel also for normal players. But the scene also got kinda lame and boring with too optimized strategies.
And in the process I think it alienated all the players for whom getting really good wasn't the goal.
I never wanted to be the best CS:S player, I wanted to hang out with the regulars on the server I was a regular at. Like a bowling league or a group bike ride.
Have watched every generation act like this since getting online in the late 1980s; act like the end of their social truisms are akin to a black hole forming in the middle of the sun.
No; it's just them realizing they have fewer days ahead than behind.
Before search engines, we had personalized search results in the form of yellow pages. The yellow pages in Chattanooga are different than the Wichita.
There was no global yellow pages where plumbers from Odessa were listed beside plumbers from Tokyo, was there?
If you were looking for something more esoteric, you were going to find different results in the World Christian Encyclopedia than you would in the Encyclopedia Britannica
That is, the www was "being killed" already in a lot of ways. "Democratization" and suchlike open ideals had already been receding for a long time.
In part, this is because of "democratization" of access. The old web was a self selected, subpopulation. After smartphones, it is the whole population.
Facebook's walled garden and other apps using the web as "merely infrastructure." Algorithic recommendations replace hyperlinks, until it no longer "a web."
Google, imo, represents a sort of intermediate stage. "Pagerank" the original tech behind their search relied on the hyperlink web structure... and also degraded it.
That was SEO paradigm. Now there is the LLM equivalent of SEO.
So... LLMs eat the web, while also making it redundant. The previous paradigm and all it's contents whill exist as a ghost within the new one.
I always hated the term "democratization" because last I checked there was never any consensus on what should be done, always MBA pricks shoving whatever they wanted down our throats.
If the timeline of things as you present them is taken as accurate, then I would say the www being killed was good, with FB opening up to everyone being the inflection point toward bad.
Webrings are nostalgic, but the old web sucked, and there was a dearth of information and fun to be had. I remember getting online and being able to confirm "nope, no new anime content worth looking at today" by quickly checking Usenet and Yahoo's index (which was updated manually) and a few webrings.
now, I'm able to watch a show that aired in Brazil the day after, with English subtitles.
People who long for the web of the 90s, I always feel like they must not be very interesting. Imagine complaining that we can't go back to libraries of the 1970s when the ones today have 3D printers, DVDs to rent, etc. because "there are too many teenagers"
That's confusing content access with monopolised algorithmic promotion/curation of belief systems, using selected news and opinion pieces to modify beliefs and behaviour.
It's nice you can view (pirate) whatever content you want, but not so nice that algorithmic platforms pretend to be neutral social spaces when in fact they're being used to promote certain political and cultural beliefs while suppressing others.
I don't know about anyone else, but anybody who goes into a "1970s library" (I don't know what that is, but I guess it means a good old-fashioned library with like, books) and complains about the lack of 3D printers or that their obscure animated series catalog isn't wide enough would embody the definition of an uninteresting person to me.
I am one of the people that long for the old internet. I don’t long for it because there are too many normal people on it now.
The old internet was interesting, open, simple, un-gated, and relatively egalitarian with few corporate interests in commanding positions. The new internet is basically a predatory environment disguised as a library.
If you judge the internet and the world by "no new anime content today" things have actually improved. But some would think that's not a healthy way to see things.
> People who long for the web of the 90s, I always feel like they must not be very interesting.
> People who long for the web of the 90s, I always feel like they must not be very interesting. Imagine complaining that we can't go back to libraries of the 1970s when the ones today have 3D printers, DVDs to rent, etc. because "there are too many teenagers"
The Internet used to be a space for high-IQ people, and now it's filled with low attention span, low impulse control people. Tik Tok and Instagram are the worst offenders in this regard, because it enabled borderline illiterate people to communicate on the Internet.
I think we still have many high-IQ people on the web. Its just they are not making tiktoks. However, they are harder to find since the pool got bigger.
The ease of use introduced the average individuals to the web and online gaming. When I was younger and online gaming was gaining traction my "clan" hosted servers. This is how I started dipping my toes into programming and server configuration. Literally, everyone that was in the clan was quite intelligent. We had debates and discussions on IRC, forums and teamspeak.
Now, gaming and the web is just trash full of brain rotted individuals. Communities have been destroyed by lack of self hosted or self owned servers. Forums have been destroyed by discord. Games are being destroyed by microtransactions.
A quick read of any old usenet arguments should instantly disabuse you of this absurd notion. Or the old archives of stupid IRC conversations.
The old internet was a space for people who were incentivized to spend significant effort dealing with technical things to connect to strangers, which has nothing to do with IQ or "intelligence" or anything like that.
There has always been plenty of people who are extremely stupid, but have no problem following technical minutia as required for things like the early internet.
Remember that early access was extremely expensive. Usually it was billed hourly, at significant rates. This means that the primary filter for internet users back then was connection to resources, through an academic connection, a corporate connection, a government connection, or a lot of personal wealth.
The quality of internet users early on, if it even existed (which I dispute), was more driven by organizations like the government, academia, and defense contractors.
And as much as people cried about "Eternal September", the internet mostly survived as a distributed place with real communities until the late 2000s, where it started to be really profitable to build walled gardens for Advertising economies.
The walls didn't really come up until the 2010s: Facebook being a dominant platform, Google owning Youtube and leveraging that to try and force their own social network, Apple controlling most phones, old forums dying, and the rise of AWS to centralize internet software design.
Being able to deal with “technical minutiae” is highly linked to intelligence. What you’re referring to is lacking social graces, which is different than lacking intelligence. Having a hobby that involved reasoning about abstract things like IRQ assignments is a pretty good proxy for IQ.
Similarly, back when computers were expensive, people on the internet generally were higher income, which is also quite significantly correlated with intelligence.
It was already on fire, AI just pointed a leafblower at the coals.
Before slop farms started using ChatGPT, they were using outsourced writers who had a tenuous (at best) grasp on the language and subject. If you searched anything on the web, your first page of search results was generally incredibly fluffy articles.
For example, you would search for "How to use python async" or whatever, and every non-stackoverflow result would go something like "Python is a useful language for async library usage. Python is a language created in... You can install Python by... Synchronous programming is... Asynchronous programming is... Python synchronous code looks like... Python async module code looks like... Popular libraries for async are... Other languages have async such as... To import async module you can..." with every paragraph interspersed with an ad. The content was surface level, relying upon blatant plagiarism of better-written articles and forum posts without attribution. It was an utterly miserable experience if you placed any value on your time.
The only thing that's changed is the quantity of slop, and the karmic irony that the content farmers who put the writing profession into a race-to-the-bottom were themselves discarded, in much the same way as the scabs who replaced striking workers were frequently discarded in the industrial era.
And that's without even getting into the subject of the people who made money as Internet point farmers of Reddit or Twitter, who would repost generic content then sell their accounts to spammers.
The Internet was already going in this direction; the only difference is the rate at which it enshittified. I agree that it's in a worse state than it was 10 years ago, but let's not look at the past through rose-tinted glasses.
Well, I called it in 2015 because they were already adding features to sites to circumvent archiving and other things. Forget AI. That has nothing to do with why the collective memory is being lost.
I find it so interesting how many comments about AI start with a disclaimer that of course it's very useful, before pointing out a negative of any sort.
I do it too so it's not a criticism, but I find it kinda interesting culturally. It's a reflection I think of the environment that we've created for ourselves, or has been created for us (or some combo) where somehow you can't point out an unqualified negative of this tech.
This to me is actually one of the clearest bubble signals, just because if it really were as good as all that, it'd be completely redundant to remind everyone of that when criticising it.
I think it might be a sort of weird collective psychological thing where we must not admit what seems pretty clear, just for fear of breaking from norms.
But, of course, yeah it's very useful for some stuff and I use it all the time :-)
Part of the problem is that there is so much overlap between GenAI's biggest boosters and people who reflexively welcome on any tech change as an unalloyed good. People can be really fanatical about this stuff. Any nuanced discussion of tradeoffs or negative externalities draws insults of being a Luddite.
My own general thoughts are that AI is extremely useful in discrete cases, will cause an immense amount of cultural and societal harm overall, and there's no clear political path to a healthier development trajectory. I don't think that's an uncommon stance, but people espousing something like that get a lot of hate in a lot of discussions.
I can easily imagine an arms control style series of agreements between the handful of leading AI companies and countries which are imperfect but still quite effective at mitigating the harms without stifling the upsides, but it seems clear that the current leading actors in this space have zero interest in that.
It's a very wide technology, there's going to be use cases where it works well. I think without the "of course it's useful" then you'd just be flooded with examples of where it works well as counterpoints whether they're relevant or not. Critics have to set the stage that they're criticizing a particular use case or aspect of the technology and not the whole thing all at once.
i think its cognitive dissonance. everyone is using it pretty frequently and so to criticize it creates this tension like "if its not actually worthwhile then why am i using it so much?" and the most prolific response to that dilemma is to write it off as saying "well of course its useful in many ways" ... with the unspoken second half of that sentence being "it's useful in the ways i use it"
It happens everywhere with every topic on the internet. You need to let your fellow tribemates know you're in the same tribe before you say anything negative, otherwise they will discard your argument and flame you to death.
Go into any subreddit and you see the exact same pattern.
I think you're describing political commentary habits becoming general writing habits.
It works with just about anyone but try and say something positive about Trump and, if you don't add a qualifier like that, the first reply is going to be something like, "I guess you like ICE murdering babies." And then the conversation you wanted to have gets completely derailed.
The majority of people's opinions on the subtopics of complex or controversial topics have to align with their opinion on the overall topic.
AI is bad so it must both be useless and be using up all the water.
Well I think theres a lot of social pressure to disclaim "yes of course its useful" because if all you do is say "hey AI has this problem," you'll have AI bros replying to you telling you about all the good and saying you're misrepresenting AI.
I agree and I think there is also a more insidious meta aspect: for every comment like this there is the worry that some X multiple more people hold that view but won't say it.
Everyone else also thinks like that and so isn't normally unambiguously critical, so you end up not knowing who really thinks that way and who's just doing it because of the game theory.
It could even be practically nobody. There's no way to tell until the deadlock breaks.
> I find it so interesting how many comments about AI start with a disclaimer that of course it's very useful, before pointing out a negative of any sort
People have to post very defensively online, otherwise some "um ackshully" guy will be immediately jumping down their throat
If you write anything negative about AI without stroking the ego of the AI bros first by praising their machine best friends they will be all over you calling you a luddite
Yes, this is really stereotypical HN: You can't really post an unqualified generalization here, because someone will inevitably jump out of the woodwork with a canned "You said X is generally Y, but here is counterexample X0 which is not Y! Your post is incorrect!"
how I handle this - if I find an interesting web page - I save to pdf & save to my external sdd. ideally I should get one of the 8tb HDD to save material I see on the web.
coz yeah good quality material is disappearing on the web fast. the only thing remaining is people building their own private search indexes.
I've been using Zotero for that, although it gets slow as hell eventually because it is a whole-ass firefox instance. They're real references so I might as well treat them like real references.
I've been working on a project that basically is a superset of Zotero, just efficient. As if you expected to have more than a few thousand references. I'm sure I'll get it done after the web is long gone.
The writing has been on the wall for quite a while, thanks to Google itself and SEO, plus Google advertising channels.
Some websites already had minimum “for humans” content for years, as it was all used to appease the algorithm.
And Social Media was the cherry on top, with information overloading, walled gardens, personalization and companies using it to pull a “Hey Fellow Kids” when using it for marketing.
Very soon, everything easily accessible on the Internet will be a never-ending loop of AI slop. Will this push us back to analogue things? Not in an apocalyptic scenario, but in general, and by analogue, I don't mean no use of computers but use of computing with a firm and predictable human touch and control.
Will there be venues left (largely) unencumbered by AI to even turn analogue? The top-down control of our industries, financial systems, education, etc by a few mega corporations and fewer mega rich people, who also have vested interests in the advent of AI as AI corps are trying to be, means there's nothing left where AI is not inserted in every vein and nerve and nerve centre.
It is as if a few people in the world are trying to turn this world into something Frankensteinian because they think that then they will get the chance to be the only ones to control this Frankensteinian. They might as well succeed. To what end? I do not believe even they know that. Blindness of greed.
Last month I rebuilt my pool's plumbing system with the help of YouTube and GPT. I was quoted $6k and it ended up costing me about $1k of materials and a few days of work. This was my first exposure to any kind of outdoor plumbing, I'd never touched PVC before.
It's unlikely I would have had the confidence to do it from YouTube alone, specific diagnostic help and a full diagram to work off of was extremely helpful. I had it prepare an SVG of the whole assembly with all the measurements and parts labeled.
I'm glad it went well for you. But I run a not-for-profit pinball museum, and I have specifically had to ban volunteer repairpeople from using the chatbots because it will confidently propose idiotically wrong solutions. Which will then get proposed to me, or just directly applied, with equal confidence. At this point I'd rather hear, "My horoscope said..." than "ChatGPT said..."
I expect that it's good for common use cases that's well documented. But then, so are a lot of other approaches.
Yeah, I would imagine pinball machines are at least an order of magnitude more complicated to maintain than above-ground pools are. That situation sounds really annoying, I feel for you.
For sure, but I think the deeper problems are how much the industry changed over the decades and the extent to which pinball repair content comes from amateurs opining on forums.
Pool technology has been more stable, and there are more people out there writing well-informed content for LLMs to extract and present as their own.
Two years ago it would have been insane to say that you got help from ChatGPT to fix your outdoor plumbing, people here would have been frothing at the mouth for merely suggesting it. Four years ago it wasn't even on the radar of future possibilities.
The quality it is today, is the worst it's ever going to be.
Maybe? It's a plausible theory. But these operations are all wildly unsustainable financially at the moment, and it's not clear where they'll get future data from, having destroyed a lot of the incentives that generated their current source content.
Bubbles are not a great time to form intuitions. WebVan [1] and Kozmo [2] also seemed to herald a new age. Decades later, brick-and-mortar grocery stores and convenience stores are still doing fine.
> The quality it is today, is the worst it's ever going to be.
Someone in 2004:
A few years ago it would have been insane to say that you got help from Google’s I’m Feeling Lucky button to fix your outdoor plumbing.
The quality it is today, is the worst it's ever going to be.
They already have all the data, all the money and all the chips. It’s actually the best it’ll ever be, as newer models will have to start paying off all that capex, newer models will be trained mostly on slop, and the SEO and influence op leeches will have begun their arms race to insert their products and values into the training data. We’ve seen this pattern before. AirBnB, Uber, social media, streaming video, et cetera didn’t get better once the VC money ran out and they needed to start turning a profit, they got much much worse.
I have schematics as early as 1937, but the odds that the manuals stayed with the machine are pretty low. (Usually these machines were owned by somebody who had a bunch, and I suspect the manuals tended to get collected centrally.) You can generally buy manuals on eBay, and some are being reprinted by people who own the original IP.
The internet has a lot of scans of wildly varying quality. The AI industry's hunger for data means there has been a lot of progress in OCR and data extraction, so I have notions of taking something like PaddleOCR and trying to turn the scans into a cohesive reference site that will work well on phones, etc. But I'm not sure when I will get to it.
If people out there are interested in working on a project like that, let me know. Email me at william@theflip.museum.
> This was my first exposure to any kind of outdoor plumbing, I'd never touched PVC before.
I find it telling that the highest praise for LLMs comes from people using it for something where they admittedly have very little domain knowledge. Domain experts usually mention major caveats. I've been testing them on subjects where I already understand the problem well, and I've yet to see any outputs that would make me trust them on things I don't already know.
When I apply LLMs to domains where I'm already an expert, I find it particularly lacking when I want it to do deep, difficult, novel work that requires precision. On the surface, the output looks pretty amazing at first. But when I turn a critical eye to every detail, I end up finding a lot of flawed "thinking", and the lengthy process of fully understanding what it generated and cleaning it up to my standards makes me question the entire value proposition. However, I find it does a great job being a low-level automaton sort of assistant.
For instance, in the domain of software engineering: I would not trust it to implement a major architectural change, or a groundbreaking, complex new feature. I would trust it more (but not completely) on something like a refactoring that may touch thousands of lines in a fairly mechanistic way, but that was a little too-complicated for simpler tools likes regexes. While that's kind of a nifty use, I think it's fair to say that non-LLM software purpose-built for such tasks can probably do the same thing more effectively for less real cost (meaning the currently-subsidized real cost of all the training and inference power burn, etc)
That's a good point - I've heard people make the same one whenever LLMs are brought up. I hope someday you're able to get more value from them.
Anyway, my pool's looking great and I gained some new skills. I probably could have gotten there with books and YouTube alone, but having another tool at my disposal made me a bit more confident.
I remember this was brought up about a month or so on HN, in the context of describing how one person's experience using LLMs can be so vastly different than another person's:
"LLMs seem good at things you are not good at."
So, if you've never touched PVC before, LLM sounds plausibly competent--it may actually be or it may not be, but you'll walk away from it thinking you learned something. If you are a professional plumber and ask an LLM the same thing, the output will more look flawed and possibly dangerous.
Same for software writing: If you're not a good software developer, you probably think an LLM is great and writes much better code faster than a human developer can. But if you are a good software developer, LLM output is slop and requires huge rework to be passable.
> So, if you've never touched PVC before, LLM sounds plausibly competent
well to be fair, it's sounds about as competent as your average homedepot employee. He's wasn't doing something super complicated, cutting and gluing PVC for an above ground pool is very common and doesn't require a plumber. I used some youtube videos to fix my dishwasher, i didn't need a professional service agent from the manufacturer I just needed some pointers.
as for software writing, for standard everyday enterprise app work which is typically just CRUD and moving data around it works fine. That kind of software does not need to be a highly tuned work of art to meet the requirements.
LLMs seem good at things you are not good at... primarily because you lack the skill to actually judge the goodness of their work, and they present their work confidently with an air of authority.
I am not a mechanic by trade, but I know nearly everything there is to know about working on an ICE car. I did paint cars professionally for a bit.
LLMs are absolutely full of garbage advice, when I try to use it for troubleshooting. However, people who don't know anything about cars are telling me it helped them fix issues. I am assuming their issues were maybe surface level and something I would just know without even looking at any manuals, because when I use it for complex problems it just doesn't work for me.
if it helped them fix issues then it helped them fix issues. Good for them. Maybe their problems didn't rise to your bar but at least they were able to get it solved. That's very useful and empowering to people without the direct knowledge and experience themselves.
That's the entire problem. The main thing LLMs gave you was confidence, and the thing to understand is that the confidence an LLM gives you is often utterly false and baseless.
Why were you not confident with literal how to videos and documentation, but became confident when a chatbot generated probable text?
Well, I guess another plausible explanation is that usually "domain experts" are people that get paid for their expertise, and have a vested interest in saying that LLMs cannot replicate what they want people to pay them for.
It’s likely it missed some crucial subtlety that will come to bite you in the ass down the line.
That’s the thing with AI - its responses sound plausible enough to non-experts but time and time again I see experts in any given field being able to identify AI content by pinpointing subtle but crucial errors. That’s one of its dangers - it gives you enough confidence to shoot yourself in the foot.
Human workers also wreck these jobs horribly that come to bite your ass in the end. I think an intelligent person equipped with AI and common sense and real stake in the thing being well done (because it's their own, so they care) is better than whatever is possible to pay for or book in a realistic timeline from another human.
Also, ask people in the trades to review each other's jobs. They will harshly criticize each other too for missing basic things and then go on to vehemently disagree. As an outsider it doesn't mean much that an expert found some fault. They always find something to nitpick.
> It’s likely it missed some crucial subtlety that will come to bite you in the ass down the line.
Maybe, but that's part of the experience of learning. I plan to maintain pools for the rest of my life, if I made an oversight which costs me down the line then the lesson will be that much more memorable.
This is an above-ground pool with a pump and a filter, the stakes are relatively low. In the absolute worst case I could rip it all out and pay a pro to do it for the price I was quoted.
A pool filter is for removing physical debris, it's not really for preventing bacterial or fungal infections.
I guess what you're talking about is sanitizer, in my case we use chlorine. I test it every time we swim, but I wouldn't have needed an LLM for that. It's very straight forward to maintain pool chlorine, my Dad taught me that when I was 13.
> It's unlikely I would have had the confidence to do it from YouTube alone, specific diagnostic help and a full diagram to work off of was extremely helpful.
People have been DIYing swimming pools for decades with the help of… books:
Still can (libraries still exist plus the internet) but finding the online resources gets more and more difficult; the fact the author mentioned youtube videos instead of someone's old-internet style "all about pvc pipes" website [0] or comic sans plumber sites [1] is already telling.
I'm not sure if the sites you linked would have been immediately very helpful in my specific case. You know this is for a pool, right?
The first is called "Everyday Uses for PVC Water Pipe" and has some cool ideas like using PVC for Wiimote holders or tridents, but I don't think that's super relevant for this project.
The second is a great forum which I'm already familiar with, but again, the topic is household plumbing and from a brief skim, none of the topics mention pools. I think mine would have been out of place.
Definitely! My understanding is everything I did has been pretty well established pool maintenance for decades now. I'm sure there are loads of good books on the topic, though I'm a good 45min from the nearest library so it wouldn't have been my first choice.
I answered something similar above, but I'll paste it here too:
This was for a largish (40,000L) above-ground pool with no hookup to my home's plumbing system. The water was all pumped in from a water truck.
The previous system was also installed by a non-professional and was mostly tubes. It leaked to all hell and looked generally redneck and awful.
First step was draining the pool and doing a nice deep clean. Then I ripped out all the original plumbing until it was just the pool outlets, pump, and filter.
I arranged it all and measured the dimensions. I fed the figures along with a tonne of photos and explanation to GPT. I spent a while talking pros/cons and landed on a design which lined up with what I'd seen on YouTube. I had it prepare me a full shopping list of PVC, tools, cements, etc. all linked to a local pool dealer. I picked it up the next day.
The PVC was all cut with a chop-saw then primed and cemented together. I found this part easier than I would have expected. I put a layer of TigerFlex hose between the PVC manifold and the pump/pool/filter inlets so it had some tolerance.
We've been swimming in it all summer, no issues whatsoever so far.
I feel like if you had wanted to you could have done it before. When I got a pool in 2020,there were still old school phpBB forums out there dedicated to DIY and maintenance. That same summer, my neighbour who in no way is techie or handy redid his plumbing too.
Yeah, absolutely. The YouTube guides have all been very helpful, a lot of them are done by actual paid professionals promoting their pool brand.
The fellow who'd done it originally is my neighbor and he's a retired teacher. There's a lot of ways to learn this stuff - my one takeaway has been that it's much easier than people make it out to be.
There are just boatloads of people who say "Oh LLMs are magically good because of democratizing access to information"
Except, for people who were gently motivated, that information was already pretty well democratized by libraries. You could trivially go to the library, and get whatever books were published on a topic, even if the only copy was on the other side of the country. Tons of the famous names from previous decades got their start teaching themselves things from a book in a library. It was very common in the technical churn of the 20th century that a new project at work meant you went to the library and grabbed books on a brand new topic and self-taught. This for example is how some programmers in the 90s developed 3D engines.
There was even a short period of human history where it was common to pay a few thousand dollars for a family encyclopedia. I got my start reading an 80s encyclopedia, focusing on the more technical tomes, before I found Wikipedia. The drive to access and learn information lead to me learning about computers in a time and place where a formal education on the subject was unavailable to me. I owe my career to it.
After the existence of Ebay, a few hundred dollars could populate a shelf with the standard reference books and material for nearly any interest. All it took was a willingness to look for books, buy them, and sit down and read them.
Similarly, the internet did the same since the 90s. Specifically, it allowed for non-physical clubs to supplement the fact that not everyone lived in Silicon Valley and could access those rich clubs on niche topics. But special interest magazines were already providing some of that functionality.
The primary filtering LLMs do is provide new access and ability to people who are far too lazy and unmotivated to do the real work necessary to learn about something without being literally spoon fed.
This helps explain why the primary thing LLMs have done is increase the noise floor of information, and explains differing sentiments. People willing to put in minimal effort to learn things had zero issue learning new information in the previous regime, so aren't that impressed when an LLM regurgitates the wikipedia intro paragraph or summarizes a popular reference. They already read that. They note that the LLMs confidence is often unwarranted, and they get reasonable results because they have foundational understanding of the domain and already know which pitfalls and problems to be concerned about, and how to prompt the LLM to make the right choices.
For people who largely are unwilling to take minimum effort to learn something new, of course LLMs feel magical, because Wikipedia level introductions to topics are magic to people who aren't already seeking them out. Of course, the question is, what in the world was previously stopping you from learning new things?
I think insecurity, over-estimating difficulty / effort, and risk aversion are all to blame. I think these things would improve a lot if people have DIY tasks as part of their upbringing or education, if only to get experience with it.
But the education system - at least for me, 20ish years ago - was very much tiered or broken up into classism: if you were highly intelligent or a good learner you'd go to advanced schools where you'd get higher level math, latin, etc. If you were "dumb" you'd get taught how to do woodworking and masonry. At best I was taught how to use a figure saw, drill press safety measures, and how to patch an inner tube (welcome to the Netherlands, this is very important. Or, was, it's much simpler and cheaper to just buy a new inner tube nowadays).
I've love to know more about how you did this. Was it the underground plumbing or your pump system above ground. This is one of the few areas of the home I have little understanding of.
This was for a largish (40,000L) above-ground pool with no hookup to my home's plumbing system. The water was all pumped in from a water truck.
The previous system was also installed by a non-professional and was mostly tubes. It leaked to all hell and looked generally redneck and awful.
First step was draining the pool and doing a nice deep clean. Then I ripped out all the original plumbing until it was just the pool outlets, pump, and filter.
I arranged it all and measured the dimensions. I fed the figures along with a tonne of photos and explanation to GPT. I spent a while talking pros/cons and landed on a design which lined up with what I'd seen on YouTube. I had it prepare me a full shopping list of PVC, tools, cements, etc. all linked to a local pool dealer. I picked it up the next day.
The PVC was all cut with a chop-saw then primed and cemented together. I found this part easier than I would have expected. I put a layer of TigerFlex hose between the PVC manifold and the pump/pool/filter inlets so it had some tolerance.
We've been swimming in it all summer, no issues whatsoever so far.
There's plenty of value from LLMs, if you treat them as an advanced search engine and auto complete machine like they are. I've used them quite heavily to research things...but these are done best in the hands of skeptical people.
The problem as usual is the mass of rich people trying to profit off of them, not the technology itself.
You don't have to let it design the architecture. Personally, when I'm using AI: I design the software, and the AI implements it. The classes and functions are typed the way I want them to be typed. AI does a great job.
Yeah if you just let it run wild, it will produce subpar results. But that is not much different from humans tbh.
When you discuss design and architecture first, and write that out in a design doc or something along those lines, it works quite well for the most part.
And 9 out of 10 times when it produces some poor results, just asking "is this really a good approach?" or just stating "This code makes me very sad" it will most of the time do a really good job of analysing why that code is bad and how to improve it.
I was talking about style and architecture. In all these cases the code is working as intended either way. There is no such thing as "correct" style and architecture. There are only tradeoffs. You should know that if you have 31 years of experience.
> There is no such thing as "correct" style and architecture. You should know that if you have 31 years of experience.
And you should know that this is not accurate. But at this point this has devolved into a dick measuring contest, which I refuse to do. Have a good day.
For well-scoped tasks I wouldn't say the code is bad, not brilliant for sure, but definitively good enough.
For a lot of problems "quick and good enough" is all that is required. I've used it a lot for managing my Home Assistant setup. Has saved me countless of hours.
In principle I could've done it myself, but I never would have, the time investment required to learn it wouldn't have been worth the value I get from it.
I’ve been programming professionally for 29 years and with the right constraints (strongly typed language, strict linting, adversarial review styleguide, test coverage, clearly defined spec or requirements) it does indeed regularly write good code. All of the other stuff required is either set-once-and-forget or is also AI, and it does require supervision, but it writes good code the vast majority of the time.
Note that I am talking about frontier models and only the last 6-12 months. Opus was really the breakthrough point for me.
I’ve been programming professionally for 31 years. With all those guardrails it writes functional code - but I would still not call it good. But hey to each their own.
I think of the distinction here is about the domains we're coding within. If your software projects are product-y CRUD apps/sites... well, LLMs will tend to perform decently at that, because it's fundamentally pretty trivial work anyways, and there are so many examples to draw from. All CRUD apps are essentially isomorphic, with just some rules to plug in about data validation, business logic, etc. On the other hand, if your software projects involve more-difficult subject matter, you're going to find it struggles a lot more.
AI can make fully functioning computing solutions. That's just a fact and denying it is like denying that a bicycle can roll and only four wheeled vehicles can.
Or saying that a car is a better vehicle than a bicycle. That is probably true, but many times for many people a bicycle is all they can get.
Glad you went for this analogy. It’s like an AI producing a board with a beach umbrella nailed on it and four wheels, and saying “meh it’s good enough for transportation, human car designers are doomed”. It’s functional, right?
It’s good at all kinds of stuff that is search adjacent, “find amesent parks with water feature within 100 miles of me that have rv parking nearby”, etc
select l.id, l.name
from locations as l
where l.type = 'amusement_park'
and exists (
select id from locations as l2
where l2.type = 'rv_parking'
and distance(l2.geo, l.geo) < $nearby_distance
)
But what we should have is a good map software where you could filter by type and distance (How many amusement parks can be in that circle?) and quickly check if there's a RV parking nearby.
I installed Ubuntu on an old laptop that I let Claude Code sysadmin. It makes it really easy to self-host open source stuff, and if there are issues it can fix them too.
Funnily enough, I've had a somewhat mixed-to-hostile response when trying to upstream the vibecoded fixes, so I suspect using an LLM to fix broken open-source software (that human maintainers don't have the time to fix themselves, nor the humility to accept an LLM-authored fix) will become more of a thing going forward too.
Oh, it's also been identifying a bunch of patterns in sales data for my business that has been increasing monthly profit consistently since last November (around $4,000 USD, every month, cumulatively so far with no sign of slowing down - could easily be $10k/mo in increased gains by end of financial year).
It can be very exhausting to receive a lot of PRs. especially when they're LLMs. We have to read what people don't often even read themselves. They're often very wordy. Then it hurts all the more when it's wrong.
I've had negative responses to small, isolated submissions that I've heavily QA'd and semi-positive responses to longer ones. I think it has less to do with the "realpolitik" of the code itself and more to do with the maintainer's viewpoints about whether AI as a whole is a positive or negative thing.
I tend to respond quite well to AI-authored or assisted PRs to my project, but to be fair we maybe only get 3-5 PRs in a good month.
I tend to value person-to-person interaction. I've received a lot of purely automated responses, which shows low value to me as a person, and ultimately the project. I'm not disagreeing that code matters, but we're losing person-to-person discourse and the community that comes with it.
Thinking it through a little bit is probably coming from having bounties on a few issues that might be contributing to my negative experience.
Perhaps the problem you're describing is actually Google's bittersweet solution to the problem of not having unlimited storage space to properly index the whole internet.
The whole text of the internet from day 1 and the index machinery for it can be stored in some number of petabytes. I assume the CIA and the Chinese state are in possession of such stores and it is a minor budget item.
Google and good is of some debate. Appearing on a list of 10 links is not a democratic representation of the web. I'm hoping this reckoning is also an opportunity for a better method of web organization
In original Google, "appearing on a list of 10 links" was peak democracy. You needed people to vote for your website by linking to it. The more votes you get, the higher you are. Sure, like every democracy, it had some unfixable fundamental flaws. But lack of democracy was not one of them.
99 % of the Internet was already slop before AI existed it was just written by hand. At least AI can give you a nuanced overview of medical information whereas before you had the exact same content about medical conditions on hundreds of SEO optimized pages. Collective memory, yeah alright.
Whatever AI is doing to how we find / share information aside, I do think that one under appreciated effect is on how software (and other things) get built. Lowering the barrier to entry for writing code and other tasks feels democratizing , but we may just not have had enough time to see how that hopefully has positive effects a decade out.
Its funny how fast we forget that when Google came out and arguably now many would laugh at the statement "all the good, companies like Google brought to the internet"
I've noticed these big tech companies use the word "democratize" when they do something that looks more like commoditizing. Flooding a market with supply consolidates their own power, by suppressing other economic actors' bargaining power and making quality controls uncompetitive.
> Lowering the barrier to entry for writing code and other tasks feels democratizing
It really isn’t doing any of that, though. When AI gets it wrong, that person needs to understand why it is wrong, and without that upfront knowledge or skill of reasoning, it’s much less direct to actually build something in the correct ways.
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
> No. The court specifically determined that the Internet Archive was guilty of unauthorized copying.
You're not wrong, but you're treating “guilty of unauthorized copying” as a statement of physical fact when in reality it just means it falls under an arbitrary rule invented by humans (namely, the law that defines unauthorized copying). This rule is ambiguous at its edges because it's not written as an algorithm or equation. It was perfectly reasonable for Kahle to believe that the rule can be interpreted in a way that it wouldn't apply and, by dragging it through the courts, have that interpretation be made the established one.
Even though the court has now established a competing interpretation, it is still not unreasonable to ask whether the law is fair and just under this interpretation. I feel that it isn't and should be changed.
Seems inevitable, doesn’t it? Expecting otherwise would have been hoping that notorious atheist Richard Dawkins somehow spared one specific god. Making websites accessible with history ignoring copyright is sort of what it does. That he would do it with books seems entirely in keeping with the philosophy.
I was initially confused what Dawkins was doing with books, until I realized that the "he" in your last sentence was Kahle, not Dawkins. Might want to edit your comment to put his name in, because otherwise you have a pronoun referring to a person named in a different comment (rather than the person named in your comment), which could get quite confusing if more people comment on the parent and their comments push yours down the page.
I agree that IA should have never done CDL but for the opposite reason: They should have never embraced DRM. Either make things available unrestricted and be prepared to defend or don't release it at all. People being unable to loan works during the pandemic may have just been the push we needed to get more people to see the ridiculous onesidedness of todays copyright laws.
My sister, a journalist, mentioned to me that she only uses google search because she had learned how to get information typically only Google indexed in the country she lives in, in a way it was not exposed on chat bots. She often has to search for information like Old govt forms released as public record with a fixed a certain format photo scanned into a pdf and indexed by Google were often on the second page of the search and beyond. But they are there. She knew how the forms looked and what bigrans and trigrams matching a certain part of form for a certain piece of information to search for and Google search has it. Like an official order on a tender notice for some government department which is no longer in the .gov.* website gave her the official's name and then she could track down who to contact in an office...
ChatGPT and other bots don't have it. Some how all these government documents became part of the government record and are the key for her to do her job.
I sincerely hope google wont stop indexing that stuff just because of a PM in search "de/re-prioritizing" ranking in a way that makes this impossible.
This is a conundrum I always found interesting. If you know how to use a search engine (i.e. knowing how to use operators and structure a search query), you're almost always able to find what you're looking for very quickly, and in most cases (well, before SEO), the results are high-quality. You'll spend the same amount of time trying to fact-check an LLM (since you'll likely skim the articles it used in generating its response _which you would have done anyway if you used the search engine directly_).
I actually took a (required) class in middle school that taught us how to use a library. Amongst other things, the librarian taught us how to use Google effectively. Everything I learned then (this was in the early 2000s) still works today, since the process of using a search engine hasn't changed very much since its inception.
So many people never learned (or never cared about learning) how to use a search engine, thus why we're here today.
Google as a search engine got much worse, especially in the last year or so. The index also got noticeably smaller. My 20+ years of Google Search experience are now failing me completely.
> you're almost always able to find what you're looking for very quickly[...] You'll spend the same amount of time trying to fact-check an LLM
I've been an expert Google user for over a decade and I can only partially agree with the first statement, and not at all with the second. Yes, a search engine alone is fantastic at finding things based on keywords if you know how to invoke it properly. However, there are lots of things one may want to find out which can't be reduced to a keyword search, because you can't have the vocabulary to search for it directly unless you already know the answer. Indirect questions such as "framework options to do x and y in z situation in this language". The best you could hope for pre-LLM was to find forum posts asking the same or a vaguely similar question and comparing a lot of options, finding out you picked a dud after spending an hour on it because it's fundamentally incompatible due to reasons, searching again, etc. It's hard to overstate what a massive improvement LLM's are for this kind of search to find and compare options for exactly what you're asking for given the context of your situation.
I guess it really comes down to how much faith you're willing to put into the answers LLMs are putting in front of you.
If you trust them blindly, then the vast experiment improvements are obvious and apparent.
If you don't, or if you're the kind of person that likes to research your sources, LLMs are a speed bump.
> because you can't have the vocabulary to search for it directly unless you already know the answer. Indirect questions such as "framework options to do x and y in z situation in this language". The best you could hope for pre-LLM was to find forum posts asking the same or a vaguely similar question and comparing a lot of options, finding out you picked a dud after spending an hour on it because it's fundamentally incompatible due to reasons, searching again, etc.
I disagree with this. Stack Overflow (pre-moderation insanity) forums and the like was and is great at finding answers to questions like this. Reddit threads were also useful for this sort of discussion. Much learning was had while reading through the comments on my way to the answer. Sometimes, doing that refined or re-aligned what I was looking for, as is common when doing research.
Again, it comes back to faith in LLMs. Sure, I can ask an LLM to give me a comprehensive overview of web serving frameworks for $LANGUAGE. It's up to the user to determine how valid the information being put in front of them is.
LLMs can also be extremely confidently incorrect. Example: I used an LLM recently in "research" mode to outline how a solution I sell stacks up to the next biggest competitor in pricing structures. It gave me a lot of (too much) information in a readily-digestible format, including, surprisingly, the "agreed-upon" price of the competitor's products per SKU.
Pricing for enterprise sales contracts is very dark arts, so I went to the sources attached to the result to confirm those numbers. Lo and behold, the prices I was given were nowhere to be found in any of those articles.
Meanwhile, I used a search engine manually to see if I could find a leaked price book using "filetype:pdf" operators, mostly for grins. Found it in 15 seconds. That still wasn't applicable to what I was looking for, as it was for an industry different from mine, but it was there.
I could have told the LLM to deep search PDFs (despite telling it to "ultrathink"), but at that point, again, what is the point of using LLMs if I can do the work myself?
Would be interesting to know if those same long tail results come up in Alt-Power [0], which also uses Google's index. So far I get what I ask for, but would be reassuring to know the whole long tail index is indeed shared and I'm not missing relevant results.
Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.
The Internet has been shrinking massively. My earliest experiences with the Internet were discovering the world of hobby OS dev around the turn of the millennium, when I chanced upon someone’s personal website talking about their OS, with source code and screenshots and dedicated forum. My mind was blown. I spent two years finding hundreds of small websites dedicated to the topic, hung out on IRC communities with other teenage OS nerds like me, and of course participated in the nascent osdev.org forum. To note that all of those websites were readily found through Google, and interlinked with their own topic webrings.
Today everything has disappeared or has been conglomerated into siloes, sanitised, focusing on engagement. You have YouTube videos about it (which is more cheap entertainment than actual education), you get some posts here once in a while, there’s Reddit where all intelligent discussion goes to die. IRC is a wasteland of idle bouncers. Then the LLMs arrived to kill what is left.
Who says the Internet is a vibrant place today mistakes flashiness with depth. It’s all empty calories, just makes you hungry for more, never satisfies.
On a whim I watched my favorite childhood movie last week, Hackers. It's goofy in some ways, but man it captures the "wild west" feeling of early and mid 90's internet. It was just you, a slow connection to anywhere, and open ports all over the place. Right after I watched it, I dusted off an old hub, connected a few external usb-to-ethernet adapters to my work PC VM's, and now run them through an OpenBSD packet filter. For no reason at all other than to feel that again: me, watching packets, having total control. Hitting a wall and having to read a manpage.
I don't really have a point I guess, other than even after being steeped in a dead internet for years (with a slow decline spanning at least a decade arguably) I need to approach what I think is the internet in a completely different way. As in, not at all besides what is absolutely required for work. We're ants in a jar now, not cowboys like we used to be.
May I suggest you to look into mesh networks? I am a huge fan of Reticulum. Using it feels like being a pioneer, the scene is very welcoming.
The pitch: it is network-agnostic. The same mesh network runs on the Internet or through LoRa radios or any other physical layer than allows the exchange of data packets. It scales from private networks to global meshes. It's the wild west. People are excited, and eager to grow further.
There are retrocommunity sites. DOS, OS/2 (eComStation, ArcaOS), Amiga (Apollo Vampire), Mavericks Forever. Amiga is quite alive community. If I am to write game, I consider this platform. Steam game will be lost in 100.000 games, and Amiga game will be noticed.
There is I2P. They have been disabling scripts for ideological reasons, so there were many websites without scripts. I have used quite exotic Charon web browser from Inferno OS in I2P. That was 15 years ago. Don't know how it's now.
There is RetroNAS and plenty of other software to enrich home network.
If I recall correctly, the film actually had Emmanuel Goldstein of 2600 and Kevin Mitnick as consultants, so despite the hollywood treatment you have moments of accuracy like https://www.youtube.com/watch?v=4U9MI0u2VIE. Probably the only time Compilers: Principles, Techniques and Tools made it to the big screen. At the yearly 2600 conference, they used to always do a big group rollerblade through NYC.
Yeah it's a whole lot better than the usual Hollywood depiction of hacking which brought us gems like creating a GUI in visual basic to trace and IP address. And more importantly it doesn't take itself too seriously which I think makes the ridiculous parts work just like much sci-fi takes liberties with the science part when required for the story.
Interesting, yeah I thought there must have been some consultation there. Especially because there's a couple instances where the hackers rely on social engineering over the phone first, and faking the "coin added" noise on the payphone. I guess a hollywood suit could have come up with that, but I doubt it
I, too, look back fondly on the early web. Lately, though, I have been thinking that one reason for its decline wasn't just corporate interests like we often talk about here. A significant number of early bloggers were middle-aged and elderly people. After decades, they simply aged out. The younger generation that replaced them (albeit not so much among OS nerds like yourself) was less likely to use a real computer and keyboard as their interface to the internet, just a smartphone. Hence long-form text died.
While we're here exchanging old man stories, I remember the web before blogs existed! I guess we can organize the internet into these eras, each of which was in some sense harder and more expensive to find available info than in the previous:
1. Pre-web. Internet is mostly about messages sent to individuals or groups. USENET organizes group discussion into browseable topic-oriented hierarchies, IRC does the same but with lists in fragmented networks. If the discussion exists at all, finding it is easy.
2. Early web. Dominated by topic focused websites, early online shops and personal home pages. Search engines suck and face strong competition from manually maintained topic-oriented directories (did anyone else here contribute to DMoz?), content discovery is mutual and webmasters help each other out by joining "web rings". DoubleClick and AdSense start to funnel small amounts of money to creators, but it's enough to offset hosting costs and in many cases can make web hosting effectively free or even yield a small profit. This encourages an explosion of website creation. Discussion moves off USENET onto phpBB forums. Every organization decides it's a cultural imperative to have a presence on the information superhighway. Finding information is easy as long as you can figure out what topic it belongs to.
3. Blogging and centralization era. The internet starts to rebuild itself around people as the primary object, not the topic or category. Directories die because websites can no longer be categorized by content. Web rings die for the same reason. IRC is replaced by instant messengers that are about connecting people with pre-existing friends, not mutual interest groups. Outside of institutional websites that exist to promote the organization, things become hard to find without highly centralized search engines because nobody is putting any effort into organizing or indexing what they write anymore: maybe you get a few tags if you're lucky. Spam, hacking and lack of SSO causes forums to centralize onto Reddit. This is the peak of the search engine era because you are forced to use Google to find anything. The power eventually corrupts the tech firms and they begin political censorship to benefit the left in 2015 [1]. Enormous amounts of information is deliberately made unfindable as part of a large-scale programme of social control.
4. Social media era. All the same problems as blogging except now the bulk of the content goes behind login walls that stop search engines from surfacing them. Video and podcasts start to matter more, both of which are unsearchable by default. Eventually video completely dominates, as few younger people want to read when they could watch instead. Firefox starts to replace IE6, and then Chrome. They bring ad blockers in their wake which starts to choke off ad revenues, so many websites from the web's first era go unmaintained and eventually offline. This is somewhat but not entirely compensated by the falling cost of web hosting. Social media remains because it puts people's faces next to everything, allowing clout farming and viral notoriety that can sometimes be monetized by becoming an influencer. The only part of the web's first era that really survives into this era is Wikipedia and Reddit, which by this time substitute monetary rewards for power tripping by a small group of ideologically driven moderators.
5. AI era. Information is so heavily scattered over so many tiny sourcelets and search engines have become sufficiently useless that full neural integration of knowledge is required, with LLMs issuing massively parallel and complex search engine queries as a backstop.
What can we predict for the AI era? Institutional websites will remain because institutions still have an interest in getting their agenda into LLMs, but visual redesign efforts will largely cease as traffic stats seen by executives show visits completely dominated by AI. There will be lots of conversations of the form, "why redesign our website to look more modern when 99% of traffic is AI which won't care?" Blogs will go the same way as the thematic websites they killed, disappearing as the authors age out. A lot of effort will be put into finding ways to block AI crawlers to create 'human only' spaces, especially by social media firms, but these will fail because AI will just be integrated directly into browsers and become unblockable - and anyway, the incentives to create will be ignored. ChatGPT style text oriented interfaces will last until inferencing capacity catches up, being eventually replaced by voice interaction and on the fly video generation for nearly all users.
Where we go from here is hard to say. Content creation was most pure in the web's first era, where people with knowledge were incentivized to share it with the world by the promise of a bit of fame combined with ad clicks to offset hosting costs. Ad blockers, social media and AI killed that world. You could however bring it back by producing a new platform that isn't like the web, one where AI and search engines are blocked via technological means (e.g. confidential computing). How much anyone would actually enjoy such a web is unclear.
Most hobby communities in my spheres of interest have been subsumed into Discord. It is yet another silo, but it does feel lively. I don't think I'm contradicting you here, I just feel less negatively about it.
I think Discord is one of the worst things to happen to the web. So much useful, interesting content gets walled off and locked away. It's not public, not searchable, not linkable, not indexable.
I'm never going to stumble across an interesting tidbit of information on Discord while browsing the web.
And even when you do have access to a particular server, you'll often struggle to find something posted a while back, even if you know exactly what you're looking for.
Closer to IRC, although they do have forum-like features. Unlike a regular forum, it has no web-facing presence (unless you rig up some kind of bot to mirror it).
I love Kagi, I am a paid user since day 1, but SmallWeb is a collection of English-speaking tech blogs, it cannot even begin to compete in diversity with the GeoCities era of the web. (to be fair: I've had Kagi employees telling me their dataset is growing larger and more diverse every day, so worth keeping an eye on it)
Sometimes I use Marginalia's "Vintage Web" search for niche topics; most results are dead blogs and old .edu personal websites that someone forgot to delete, still a vanishing minority of anything one could find in 2001.
Well, just as the "internet" supplanted newspapers, magazines (gosh those classic gaming mags), and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
Why should I have to suffer through Fandom raping me with screen-obscuring banners and "PLEASE ALLOW ADS" just to make some sense of fucking Warhammer 40K lore (written by unpaid volunteers anyway)? instead of just asking ChatGPT what the fuck Globriznaroks is/are.
Why should we support shady companies by sitting through their ads on YouTube videos for minute topics instead of just asking AI for the shit I want to know about?
Why should we submit to the whims of 3 mods on a subreddit deciding what thousands should get to see (fuck /r/AskScience) and then getting low-effort answers or outright trolling anyway? instead of just asking AI?
This is better for you as an individual, but worse for society as a whole. As you've noticed, monetization is the weak spot. AI will not escape being ruined by monetization, but it might be harder to notice when it arrives.
Fandom and Reddit are not the old Internet, they’re ad-driven middlemen that create nothing themselves. Self-hosted or at least self-maintained sites were the old Internet. After that point, the rot had already set in then, it just took until LLMs for it to metastasise.
100 half-baked sites hosted on Geocities, Yahoo, about pointless stuff, covered with gif-vomit that looked like epilepsy simulators?
Serious question: What do the rose-tinted glass wearers actually think was of objective substance on the old internet that's nowhere to be found now?
You can find random pointless stuff now too, just that except Geocities/Yahoo it's Intsagram/TikTok/Twitter etc.
If you mean self-hosted websites, they're still here.
If you loved all the Flash toons on Newgrounds etc there's unironically a lot more shorts and animations on YouTube now, if you but search for them (I suggest Weebl, David Firth, Sechi, to start with, and let the algorithm soak up the weirdness)
Popular wikis such as Minecraft and Runescape have successfully migrated away from Fandom. Note that Fandom leaves behind the outdated old wiki and refuses to allow deletion because it drives their revenue. After a few years, the Fandom wiki is severely out of date and loses traffic.
What we actually need is a browser that filters out bullshit.
Duckduckgo has been incredibly worse with results for me for the past year. I had used Duckduckgo for about a decade, and now it is just littered with AI generated content for search results. I recently switched to a search engine that uses Google results.
That's 100% correct. Since Google's "helpful content update" Webmasters get massive amounts of "Crawled, not indexed" reports for anything Google considers "more of the same" or "thin content". If you're not an authority on a subject, simply meaning: you already rank for similar content, or if you don't get links from more popular domains, your content is in the abyss.
It's all under the guise of "We're fighting SPAM", but the algorithm (or model) they use is heavily skewed towards intents (actions) and brands (because they 'trust' big names).
And it's not working.
A simple, short informative blog about a tool you used that could be of interest to max. 100 people on this planet is no longer getting ranked, if it gets indexed at all.
Those posts tick all boxes: no incoming links, no authority, thin content.
It changes somewhat between "Google Updates", but it's pretty clear that it's no longer working.
Multiply the 100 people not finding that post by millions of queries and it's now a big problem for Google.
My dad asked me the other day if Google got worse because smaller companies are not paying Google enough money.
He didn't see any difference between ads and content, because all results are now big brands only.
"Helpful content" is such a misnomer. It removes all helpful content in favour of AI overviews and only shows intent-driven, commercial content.
We're watching the end of Google's hegemony for sure.
This is fascinating. I haven't really kept up with SEO for years, but was recently helping an academic publisher setup their web presence and ran into exactly this issue: Google was/is refusing to index the actual journal articles on the new site and they were moved into this liminal "Crawled, not indexed" state before vanishing entirely. Comically, it's now easier to get your content indexed and served with a visible link by ChatGPT than it is Google...
I can relate to this so much. My interests tend to be pretty niche and I have little interest in most of the major online websites. The scale at which Google has hollowed out the non-corporate internet boggles the mind.
The earlier web era felt like this unimaginable realm of freedom and exploration to me. I stumbled on countless novel sites that were interesting, helpful, and/or entertaining. That content has decreased by orders of magnitude since then. These days most searches return a full page of SEO slop that all summarize (badly) the same source from years before. There's usually zero new information, personal touches, or community attached.
> We're watching the end of Google's hegemony for sure.
Who is standing by to replace them though? OpenAI and Anthropic certainly not, they are burning money in a fire pit to stay alive. There is no way in hell they can afford the compute necessary to replace Google.
If ChatGPT goes bust, they'll find another ChatGPT, not another Google. I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
> I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
Tiktok is cheap to run. AI however, it needs absurd amounts of power, RAM and GPU compute capacity to run... and there's serious constraints everywhere.
Why do people insist on making predictions for the future with today's numbers and constraints? Inference costs are plummeting, mostly driven by Chinese inventions.
Try Marginalia. It's a breath of fresh air. It probably won't find what you're looking for because whatever you're looking for probably doesn't exist in the small web.
About 15 years ago, I won a phone in a contest. Last week I tried to find information about it, but I couldn't. No AI, nor Google could find anything the contest I won it in. When people say "the internet is forever" that can certainly be true, but it isn't for everything.
That reminds me - my wife had a friend who was killed by a shark (no joke). There were newspaper articles, but Google has completely forgotten about this and the newspapers are often behind a paywall, changed their URL's or simply the article vanished from those sites as well, which certainly doesn't help.
I wonder how libraries, who have traditionally been the ones to archive the news, have kept up with everything moving online and now being subscription-holed. Hopefully it's not just the Internet Archive doing this, which has its own problems.
My library system has an aggressive book weeding policy where books that haven't circulated enough in about 2 years get binned. the excuse is other libraries have a copy available for inter library loan, but of course, those libraries have to weed as well.
so no they are not archiving their own core books let alone periodicals
The only way to be immune is make ads and tracking impossible or at least not profitable enough for bigcorps to care about putting them there. So, maybe a web without Javascript, images, or videos, lol. Something truly primitive like the Gemini protocol.
Kagi can’t solve the fact that the internet itself has gone to shit. The independent forums are largely gone, the mainstream social media’s have all locked down and been spammed with AI slop.
The combination of mass SEO spam/slop and Google's intentional lobotomization of search has lead to it being next to useless. I also strongly suspect Google censors topics at the whims of various government agencies (this was very obvious during COVID and leading up to the 2024 election).
They're so horrible that I've started defaulting to their AI summaries. And I hate those summaries. It's just that the regular results are so terrible now, and seemingly getting worse at a noticeable pace.
I used to not worry. I was sure that a competitor would come along and fix search. But the longer that's not happening, the more nervous I'm getting that we'll actually lose search. If a few more years pass in the current state, I'm afraid the majority of people will forget what search was like and default to AI summaries.
I've tried alternatives, including Kagi (not actually relevant because there's no way I'm – directly or indirectly – buying Russian products) and Uruky, but they're not good enough.
(Edit: Added "directly or indirectly" about Kagi to point out that I'm not claiming that Kagi itself is Russian.)
Unsurprisingly all search engines seem to be struggling with AI content sites as well. It's rare to get human articles, sometimes rare to even get authorative websites. It's frequently a bot site with a plausible enough name like, potterspainterly.com or medhealthdirect or something with oddly specific articles written in the last year.
This is the root problem. The idea that Google is deliberately sabotaging search seems far less likely than the idea that the internet is mostly garbage and SEO slop.
Google used to penalize websites for SEO hacks. Then they started attending SEO conferences themselves. SEO specialists didn't outsmart Google, Google stopped trying. As the sibling implies this is most likely because they noticed that the easiest way for SEO spam to monetize is ... Google ads.
I stopped paying attention to Brave years ago when it started fiddling with advertising-linked cryptocurrency and content injection. Is there any reason I should extend it any trust now?
Thank you for your suggestion. I think I've discounted Brave automatically because my brain is numb to the dime-a-dozen chromium browsers out there. I'll definitely give the search a try!
If I recall correctly Brave scrapes the web via their users, cannot be individually disallowed in robots.txt and Brandon Eich is conversing in a pretty hostile manner in every thread about him or his company.
I'm not in the brave ecosystem otherwise nor a big fan. But the search engine was competitive with google when it was still cliqz, before it was shut down there and the leftovers bought by brave. And it still works really well.
Even if brave were problematic it would be the lesser evil to me.
Probably what is happening here is that in the race for AI, which, whether we like it or not, means power and control, Google crafted things in a way that it does not do "self-competition" by their old search engine. Idk, just throwing ideas aloud here.
That's certainly the goal, to drive more users to AI by making the other option worse. The real question is when do you put a stop to it. When brain implants are 2x as productive are you going to say nah? It's times like these were you are supposed to take a step back and consider what you are actually producing and why.
In my case: Russian state is my direct enemy and most likely to invade my country. And would do it if they would consider success likely.
Previous wars with Russia were obnoxious with very bad consequences, so I dislike idea of even very indirectly funding them.
And I support actions that are harmful to Russian economy, also when they are harmful to me - as long as it is not too badly balanced. As this is much cheaper than directly participating in war.
(I am from Poland)
PS
Yes, I understand that at some point there are some indirect effects that you cannot avoid.
I also understand if for some people paying Kagi that pays tiny fraction of that to Yandex that is paying taxes in Russia which funds their wars is too tenuous connection to care.
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
Maybe the curl thing is because they are happy to let you do some light scraping. What they want to avoid is bots directly crawling the page interactively. No one seems to be blocking chatGPT when I promot it to use it's web search skill anyway.
Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
I think the other thing they're trying to do is get most of the internet to send them all of their cleartext traffic. Expect in 2040 the PRISM2 docs will get leaked by some Eduardo Rainedon and we'll find out Cloudflare was the NSA all along.
It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.
The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up...
How I can strugle to keep them off? To peer to network, you need to talk to human.
Arrange L2 connection, assign IPs (only static). If its leaf node, we are done.
If its another network, we need to form BGP connections to exchange routing.
Yeah, Network by Humans for Humans. Thats why Im not interested in all those IoT/Auto networks when you just connect and stuff automagically configure. It looks nice at first glance, but you loose control. F2F works way better in that matter, like RetroShare, but I never investigated it much.
Nah, it needs to be IP. IP is well estabilished protocol, everything speak it.
Once you set it up, you can use it whatever you like. Web pages, gaming service, IRC, Mail, P2P confereces, everything. Everyone will bring it own slice to the pie. You love networking, became PoP and peer and provide access. You just want content? Connect to closest PoP, get IP + DNSproxy and vioala.
I know DN42, but this network is more oriented toward R&D and experimenting. Yeah, its not for your avarage Joe. But your avarage Joe can buy connection from VPN provider and use it, and so we can provide user friendly PoPs with minimal skills needed to setup, supporting different VPN software.
I just used a VPN yesterday and the NY times blocked me because they think I look like a bot. It gave a couple possible reasons, one being "a bot was also using this IP address".
VPNs are great for torrenting but any serious website like an online bank or web email provider will turn you away. They claim it's for bots but really it because they only want customers they can track.
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
Those hosting content for ad views, or for fame or other personal gain.
Sites like Wikipedia or developer documentation pages which exist to distribute knowledge for its own sake don't have any reason to care whether that knowledge is consumed by a human or a computer being used by a human.
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
It knows you, but it doesn't know me. Perhaps you are legitimately noteworthy enough to not worry about this!
My experience being on the searching end is that these things are terrible about attribution of where they find anything. Which has bad consequences not just for authorship, but for correctness (which is the usual reason I'm poking at them -- they're being wrong again). This makes a lot of sense when you consider the massive, massive compression that's got to occur during training, but it's still frustrating.
There's massive differences between "my site isn't hugely popular, but I get some readers", "my site is up and findable but genuinely no one visits except scraping bots" and "it's available, people would visit if they knew it existed, but they're not being offered it, they're being offered bot distillations with no reference back".
The last one is the worry.
The middle is... where we all start.
The first is not a bad place to be, all things considered!
So I'll write a lot of falsehoods to poison the AI. Like the urban legend where the sky is blue. I'll say the sky is blue, AI will think it's true, and regurgitate it to unsuspecting users who will see that AI is completely unreliable.
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
> "humans would visit the website and the creator would get some reward"
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves.
AI will mutate the knowledge and eventually produce some mangled version of it, mixed in with ramblings from a random reddit post and half a paragraph from a copyrighted book that its fascist creators stole.
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
> This should be "This should be 'You're thinking'." don't you think?
Reflexively, I think it should be more like ...
javascript: `This should be "You're thinking".` ;
// to preserve the original character use and to avoid '...'...' parse foos
// however `"...".` also possibly deserves a [sic] to critique the original
// i.e. ~grammar police say the period belongs within the quote marks, no?
I finally got the downvote powers yesterday, and I was irrationally pleased about that.
It's not even about votes, I don't even need votes. Whenever I write something that I am happy with, I read and reread it imagining I am reading it as a third person. Sometimes it forces me to rework my arguments. I wouldn't write to convey my ideas through a chatbot. And that's what this post is about-- killing the internet and with it decimating any audience you might have accrued if you had something to say and you published it on a website.
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired.
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
I had a flippant answer which was “who ever wanted to write for a machine in the past thousands of years”
But maybe that is the future.
Every country starts erecting their own towers of babel that we talk at, and it constantly compresses our conversations down to the most effective distribution of weights.
At some point talking at the machine becomes a high status job, and we give respect to the people who whisper to it the most.
Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
No, cheap books came from the invention of cheap printing. But the quality of works was going down because printers were just printing with zero copyright protections. Copyright is what brought the wealth of works worth reading that were then printed using cheap printing. If the only money is in private works for private libraries, that is where the quality stuff is going to go.
Making the avenue of creating for the average person also the avenue for the most income was huge in creating our modern literature landscape.
It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there.
Disagree. The printing machine was killing the industry allowing copies where people earned nothing. Copyright was created to ensure quality works existed to be copied.
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
> Funny, I published information in the hopes that humans would benefit from it.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
That works now because there are human made sources that the AI can find and summarize for you. But now there are no incentives at all for humans to write anything on the internet and if they do the content will be buried by hallucinated content someone else posted at a larger scale.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
Isn't this only a transitionary problem though? Right now, during the transition there is no incentive for humans to write anything, it will get drowned in AI slop.
As time goes on, more and more people will recognize this problem and we'll develop new ways of measuring information quality and trustworthiness. Nothing about this problem is fundamental, it's just that we're in the middle of a very chaotic transition.
AI seems to be on the same trajectory? Search was very useful in the start also, until it became entrenched. Then search placement became a target, and they are just focusing on extracting rents. All way paying the content providers zero or near-zero. And with years of that dynamic, we end up where we are now. It was the same with "social media". The same will happen with AI. AI is a power for more enshittification - being currently less shit than Google is (mosy likely) temporary.
>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I don’t agree with the parent that the solution for SEO spam is paying for the search engine (I think there are other good reasons to do this though), especially since people have been doing things like naming their company “AAA Auto Repair” to be first in the phone book since before computers ever existed. But the person you originally replied to does have a point in blaming Google for the problem. Most SEO spam sites make their money from ads, and Google are the ones who run the ad network, which means Google are the ones funding them and creating an incentive for them to exist.
> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
Ah, the argument is that Google intentionally avoids showing you the most relevant results, or at least avoids "solving" the problem of webmasters who attempt to 'game' the system.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I'd frame it more that Google is incentivised to show you the results which are most lucrative for them to display, rather than the ones which are most beneficial for you to see.
Incentives. Which is why platforms (such as search) should never be allowed to be in the same company with things built on top of them (such as ads), if the combined company is significant for an important market.
We used to know better, Standard Oil vertical integration was dismantled.
That only works when there's an abundance of documentation for your specific device. The moment you're on a more recent version of something and have a weird issue, all hope is lots. You get stuck in loops because all LLMs keep recycling old advice that no longer applies. This problem will only get worse and worse as people no longer as questions on public forums, so answers are not publicly available either.
While I understand that you feel good as you got the device configured, you would have been better off without gemini.
The 'old' way as you call it would have resulted in you knowing the backgrounds and the inner workings of your router, making maintaining it a breeze and helps you actually understand your setup.
Besides that, it would probably trigger you to rethink some of the things you now blindly have implemented because gemini did not show alternatives nor reasoning behind it. (something that most definitely would have been documented on the source pages)
I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
What on earth LLM are you using and do you tell it which distro/version of Linux they are supposed to be working with + give them access to web search or man pages? I haven't had this issue in like a year and a half.
By default I use Leo, and I do specify the distro. I tried to qualify by version, but when I do that I lose out on a lot of correct information for things that haven't changed in awhile.
Leo is not even the LLM, its just a browser extension typically running a very weak/cheap model. Why not use claude code so it knows your system without relying on you to give it accurate info? I guarantee your results will be night and day
I've had the opposite experience giving claude code SSH/ADB access to my devices. Not sure if it's the model itself or the harness, but I haven't had to manually do a sysadmin task in months.
Yep - it's awesome, the problem is that Google isn't sharing the revenue with the context creators anymore - over time, unless fixed, this will decimate the knowledge base it feeds on.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
Yeah I mooch gemini off of google but they've def made a few cents from me via ads. Also apparently I'm a rarity because I don't run ad blocking (except firefox itself does a bit according to a few websites). Its a very good deal for me though.
I hate to break it to you, but we are at the start of enshittification circle. Once Gemini has monopoly you will reminisce with joy google search results.
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.
Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
Thats probably the kicker. I think about this some, also being myself a gemini chat mooch. How good will gemini be in a years time? Its great now but will there be an incentive at some point for it to become shit? Free tier is awesome now, but is it a loss leader? Training on my chats, maybe in a year my chats won't be worth the free access?
AI is on the whole bad but it’s impossible to forget that the ostensibly human-curated Web as presented by modern search engines is bad too. Invasive ads, popups, and worst of all the substantive content has a nine inch frame of SEO filler and a three inch picture (what you were after). And some things are not even high-tech slop stolen. It is just old-school verbatim copied from another website and repackaged with another frame.
But yeah, text remix machines are not a long-term solution to that problem.
Even with the /s I think you misunderstand. If you just want to configure your router, then AI is neat. If you want to learn and understand what's going on in the router as you set it up, then being given the answers basically teaches you nothing.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
Why would most people want to learn how a router works? It's the end goal that the vast majority of users want - i don't want to learn about routing tables, I just want to expose a port for plex, etc. Also, I could have AI craft a fascinating and engaging way of teaching router setup that actually helps me learn instead of wasting time searching through random forums form google results
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
Personally it's because i used it when it was brand new and work paid for it and i had no idea what model it was using, my boss just turned it on for automatic PR summaries and code reviews and it was universally dogshit, and the auto complete in my IDE was awful as well.
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
Interesting, I've seen much better results on DDG. Most recently was the search: `site:feeds.bbci.co.uk inurl:rss.xml` which works on DDG but gives zero results on Google. As far as I can tell, Google just decided not to index these.
Yeah I mostly still use Google out of habit but there have been a few times where Google has decided something isn't worth indexing (too niche, doesn't use SSL).
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
I find DDG is struggling to, or chosen not to, filter or derank obvious AI generated content farm sites. Of which there are an insane amount of already.
We thought blogspam was bad, at least it was easy to ignore. It's hard to find authoritative sources for a number of topics, worryingly health advice is one of them.
Idk, Google seems worse there too. I tried "my left hip is hurting".
Google shows an AI overview and "people also ask" with zero search results above the fold. If I page down I see a single search result for clevelandclinic.org, followed by youtube videos and image search results. The next page has a single search result from rush.edu and then "discussions and forums" which has Mayo Clinic and Quora.
DDG also starts with the AI overview (although I disable that) and has two results from webmd.com with deep links to multiple pages on the site, all above the fold. Then the same clevelandclinic.org result as Google but again, adding deep links to other related pages.
I can't comment on the quality of webmd, clevelandclinic, or rush, but Google pushing the user to youtube and quora for medical advise seems worrying.
The problem and what the article is pointing out is that original content is slowly and progressively being replaced by AI content. And since AI gets trained on this content as well it will eventually train itself on previous gen. content that was also AI generated. It is slowly eating the web. Eventually you won't even be able to evade it because it'll be everywhere. Before AI became this expressive I could at least expect someone writing articles, FAQs, blog posts to have some backbone. Now I frequently run into content that obviously was never even checked by a human.
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
It seemed like Google was good ten+ years ago and then gradually trended towards being rotten around 2020-2022. In those days blogspam was king. I want a cooking recipe and it would show me a life story. Or I wanted a OG web game but a link farm would come up top.
I think there are leaked internal comms where they discuss nerfing search results to pump engagement and ads impressions. No doubt it would boost ad bids too if businesses couldn't be found organically.
Around that time Bing and DDG were actually better. Then LLM's came along and they started to take things seriously again.
Maybe they think the OpenAI threat has abated enough to begin enshitification cycle 2.0.
I know Google started deteriorating. It was in 2008 or 2009. Up until then Google would return 0 results if it couldn't find a document with all the words you specified.
Then it started serving synonyms, attempted to correct spelling, and so forth. Instead of serving up that there was 0 results, it attempted to be "helpful".
For a while, you could enable "verbatim" search, but even that has gotten corrupted.
Their search quality has deteriorated ever since that change. It is sad.
I've used DDG for years (thousands of searches) and when I switch back to Google thinking I might be missing something... I'm always let down. Seriously, the results are pathetically bad now and have been for years.
that control (so I can turn it off completely) is why I picked DDG to replace Google, who force feeds us the hallucinations
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
Yeah, there are times where DDG has like, literally three results. Yet, I know for an absolute fact, there are hundreds of pages on the web that contain the terms I specified. Web search is becoming utter garbage.
I don't think it's really an "AI problem", we just got to the "worse" part of the "Worse is better".
Back in the day one of competitors of the World Wide Web was Project Xanadu. Project Xanadu was supposed to address the concerns like content persistence and version management within the core design. As such it was much more complex, opinionated and centralized.
WWW on the other hand comes with no guarantees - you might get a document in response to a HTTP request, and that's it. But WWW service can be rolled out in a completely permissionless way, and is quite simple - effectively, the contents of the file system can be shared with the world, so e.g. a document can be published just by putting its file into a particular directory within the
file system.
Thus Web could get to a "good enough" state much faster and quickly spread all over the world. But its permissionlessness and simplicity lead to downsides: impersistence and chaos of broken links, web search provided by mega-corporations, etc.
WWW evolution was, unfortunately, not "incentive compatible" with features like advanced persistence and identification clarity: there was much more focus on entertainment content and ads
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
Think of the more malicious use-cases though: The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.
Gemini has been a hilarious companion to my while I fixed the balance shaft chain guides in my old Mitsubishi triton (mighty max for US readers).
First it told me I could just remove said balance shaft chain as an emergency repair. Sorry Gemini, it also drives the oil pump.
Then it told me I could remove the water contaminated oil caused by removing the timing case by filling the crankcase with hot, soapy water and running the engine. Lord no.
Then it gave the wrong instructions for putting new gears on the balance shafts which meant the chain guides didn’t align with the chain. I’ll do it my way thanks Gemini.
The rest of the mistakes are too trivial to recount and sure it’s a pretty obscure subject but if I trusted it with a topic I’m not familiar with there is a huge potential for damage if you blindly follow it’s overconfidence. I miss normal searching.
Meanwhile, ChatGPT correctly diagnosed what was wrong with my plant from a single photo, identified which leaves I should cut, and annotated the picture showing where to cut and what not to touch.
I honestly expected a made-up useless generated image that matched the idea but not the actual thing.
> While the web has always been organized around intermediaries that shape what survives online and who sees it,
This statement, from the sixth paragraph of the article, is something that I would have liked to see addressed more in the article. The article implies that this is something that must always be true, or cannot be changed, and simply focuses on how we could have better/better funded/better protected intermediaries (AKA gatekeepers), and doesn't discuss the possibility of an internet (or part of the internet) without gatekeepers (and doesn't ask if it has ever existed/does exist/should exist)
I think relatively few people have fully comprehended the enormous downsides that LLMs bring with them. I do think they are useful tools in many situations, but the fact alone that they have "flipped" the "takes time and energy to write something useful so others can more quickly read and comprehend a complex idea" equation is very, very bad.
Not just reading, videos have been ruined in many ways as well. My base reaction to seeing anything surprising on video form has switched from "Interesting!" to "Probably fake". Novel events caught on video are now immediately suspected as being AI and discounted by many people.
Same for software for that matter. GenAI has killed my interest in releasing web apps outside of work. Before I used to find value in putting stuff on the web for others to find and enjoy, and it sparked some interesting interaction with other people. Now anything I publish will mostly be consumed by bots and thrown into an AI blender that completely divorces it from the creator.
Indeed, I also don't listen to music after 2024. I'm also not interested in anything my friends have made (with claude or others). Or how so much of creativity is now a prompt away.
A lot of our culture is disappearing before our eyes and people are defending it.
And make sure "follow" means "buy their stuff on Bandcamp or directly from them". Streaming platforms send most of the money you pay to bands you probably never listen to.
You're comparing almost three millennia of text, to a couple of years worth of text.
Let's see this play out.
Plato also threw shade on writing. Which is an invention that turned out alright for humanity, imho.
Maybe it'll turn out that we can push through to a yet-unknown-but-better state, as we've so far done, rather than relying on just going back to a previous state that we were content with.
Sorry for the somewhat sarcastic tone. But the hyperbolism deserved it.
The comment they’re replying to didn’t add anything new - it’s just the same old regurgitated ideas and has a lot of misconceptions. Almost exactly like LLM content. Taking the time to respond in detail would be a waste.
Nothing they said demonstrated any understanding of what they replied to. They simply restated their opinion with no reference to any of the points made in the comment they were ostensibly replying to.
Am I really the only one who doesn't care whether or not prose was written by AI? Assuming that a text is signed off on by a person, then surely what counts is its substance, not who or what actually strung the words together. I'm constantly surprised by this obsessive need to falsify what is unfalsifiable. And among nerds of all people. It feels religious, like a modern form of heresy-purging.
I care because in most (not all) cases, (a) AI UGC is one- to few-shotted with no concern for accuracy or readability, and (b) why would I read the results of someone else's prompt if I can achieve more or less the same result for free on ChatGPT?
Now, this is for disc brakes, so not applicable in my situation. I've never ridden a bike with disc brakes, so I can't vet for the accuracy of this guide. Maybe it's totally right; I'm sure someone reading this will know.
However, just look at this article! Zero pictures for a process that _really_ needs it. It just drones on. Imagine following this guide only to discover that the 12-15 Nm rotor torque the article is asking for is actually too high!
A similar article from before AI would at least have (stolen) pictures describing the process. It would also very likely be shorter. Which creates the rub: if I know how to write something engaging and correcting AI's copy will take _just as much time_ as writing it myself, why would I use AI to do it?
In short, AI content is lazy content. My time is valuable and finite; if I'm reading your stuff, I want to know that you at least tried to give a shit about my time as a reader while producing it.
> However, just look at this article! Zero pictures for a process that _really_ needs it. It just drones on.
Buckle up, because the next phase of this AI content generation nightmare will be articles like this featuring AI generated images of a bike repair in progress, except the bike will only kinda-sorta be like the real bike the article is describing.
Sure, I get that argument. I just don't buy it. What counts is whether the article helps you fix your brakes. If it doesn't, it seems immaterial to me whether or not the bylined author sweated over it for hours. Not least because I have no means of knowing.
But the OP already addressed this-he can just ask ChatGPT directly how to fix the brakes. He’s gone looking for a human expert because (for whatever reason) he doesn’t want that.
For me the issue is very clear: I happily use AI to answer tons of questions each day, but if I’m reading your website/blog/article/Jira ticket, I expect real human input.
It should matter whether the author “sweated over it for hours”, because that means they actually considered the best way to communicate something. If they didn’t do that, then they clearly don’t care enough about about the quality of their output and I shouldn’t spend any time at all on it. In fact it’s worse than that because if the author doesn’t respect the reader enough to craft their output, then I as a reader have no respect for their work, and I resent that they’ve taken my attention for the time it takes me to realize it’s ai generated. If you’re publishing Ai content, then I can only presume you have an ulterior motive than plain communication- in the case of OP I’d assume this is for ad revenue or clicks.
There are many types of writing, and I struggle to think of many where a statistical model could feasibly produce an acceptable output for the given intention.
Eg. A poet chooses their words extremely carefully; a good instruction manual is written by the designers of the product (not just guessing at a common or plausible method); a news article has an angle/story beyond just the facts.
If you or OP suspect (suspect) that a blogger's taking shortcuts, and that's a problem for you or OP, then you or OP can just skip that blog and find another that appears (appears) to meet some criteria for labor input. To the extent that this is all unfalsifiable, it just should not matter.
I'll go further. Even if it were falsifiable, why should it matter? An example: I subscribe to a reputable US publication and I enjoy its journalism. If it turned out that one of its writers had (somehow) used AI to generate their article (which I enjoyed) in a click, here's what I would say to them: Hats off to you! How did you do it?!
I don't care if the article took them two minutes or if they used AI any more than I care if they used a spellchecker or wrote it while standing upside down. Why should I? The author put their name to the article and I got something out of it. That is all I was ever looking for.
> Assuming that a text is signed off on by a person, then surely what counts is its substance
Unfortunately that's a big assumption and for many their "signing off" process will amount to "skimmed it and seems ok?".
When I get an LLM-generated doc or runbook, my first thought is that its very possible that I'm the first person who has ever read this. It used to be that writing something took a big time and energy investment up front by one person so that many others could comprehend the ideas with less time and energy needed. LLMs flip this equation around which is a very bad thing.
The problem is that AI writing uses a pretty distinctive style and when you read a whole newspaper that was written or edited by AI it is very monotonous. Also, I can't really trust it because I know that the author probably phoned in his fact-checking.
Yeah, I keep feeling the same thing once I know I'm reading something or watching something purely AI-generated - that there are good points being made, but the delivery is just so off that it kind of ruins the points themselves.
If people used AI to get the scaffolding of their article going and then re-humanized the entire article, it might be really good, still at a fraction of the time it would have taken to write it themselves, but no one seems to want to expend enough effort to go the extra mile.
With video it's pretty well hopeless, unless the AI helps accelerate fully human traditional production techniques.
With music, there might be a middle ground, I have had some success getting AI to generate good sounding loops to use in electronic music production, but it was just a matter of brute forcing enough outputs to finally get something decent, which isn't fun at all.
Why shouldn't people be upset if authors are being replaced by like a handful of the same ghostwriters?
An LLM can't convey your ideas and your voice better than you can convey them to the LLM, right? It's a middleman between you and those you want to reach that adds more points of failure, an additional entity that must be understood and made to understand, or else meaning is lost.
EDIT: Never mind the motivations behind a majority of LLM-generated articles, which is to be One Unit Whole Content, a colloid for ads.
I would be fine with it conceptually, if the result was not as grating as Claude's ramblings about load-bearing seams and the real shape of the problem.
It's just exhausting to read, and it displays a lack of effort put into communication.
people who generate entire blogposts with claude probably do not care about substance, tho. before llms, people would have to put effort into what they write and be mindful about the things that they are saying. now it just generates the whole thing, the people who "signs it" barely reviews it fully most of the time, IF they review it at all, proceed to say "eh, good enough" and publishes it.
Ironic anecdote. Although I don't care about AI, I'm triggered by your own carefree attitude to orthography (no capitalization, unpronounceable words like "llms"). At least AI is easy to parse!
The former. As did everyone else who read it, including you. Capitalization and punctuation are not decorative extras, they were invented for a reason.
In addition to the sibling comment, which I agree with [edit: the one starting with “people who generate entire blogposts with claude probably do not care about substance, tho”], I want to add another point. Even when the substance is key, the delivery has a quality to it that encodes a useful signal.
To illustrate what I mean, imagine two blog entries on the same subject, but one is written by an expert on the subject and one by a layman. As humans, we can generally (not perfectly) tell which is which. The expert will write in a certain quality that comes with expertise and is genuinely difficult to fake without it.
LLMs disrupt this pattern. LLMs are able to write with the expert’s quality even while writing bullshit. I don’t mean quality as “measure of goodness” here but merely a set of traits or properties. Humans reading LLM text are far more likely to misjudge the author’s level of expertise.
The next step is unscrupulous humans exploiting this to trick unsuspecting readers into misjudging a text, and now you have an internet where you can’t trust anything anymore, at least not at first glance. Even relatively discerning readers now have to waste time reading more of the text before being able to dismiss it as having substance of little value.
So, as I see it, the upset is not simply from AI provenance of a text, but from this level of dishonesty and subterfuge, coupled with the helplessness with which we’re exposed to it.
Yep. Almost all of the top results I get from my searches are AI-generated garbage if the content was published after 2025. It's essentially guaranteed now. Whatever; let the people have what they asked for. At least there's an infinite supply of old books.
I was just thinking yesterday that it might actually be worse than nuclear weapons. Both in terms of likelihood to destroy all of humanity and in terms of “how did nobody working on this see that this was an obviously horrifying idea”
Come on you know what he means. Blog posts, articles, that sort of thing. AI has definitely made reading things posted to HN a much worse experience (even if they aren't all slop). It doesn't seem to have infected the comments yet, mercifully. I guess for a quick comment it's still easier to write it yourself than get Claude to do it.
It has infected the comments. Our patrons are fighting a good fight, but good slop is mostly indistinguishable from a worthless karma-farming comment. Mods are probably just cheating because there are people with 15-20 year old accounts, many of whom they personally know, and you can assume that the stuff they're interacting with is not slop or is at least worthy slop.
What's more realistic; 50+ year olds of today just parroting sensory experience where they heard their dead or dying elders complain about the kids back in the 00s, 90s, 80s, 70s, 60s... etc etc
Or the 50+ year olds actually figured out how everything must work for the next 1,000 years they won't be around for
The olds who grew up in a PTSD addled post world war and cold war social reality while huffing leaded gas smog? They figured it out forever, everyone!
...No. You figured out yourselves relative to technology of your day. Tech will change and the living will figure themselves out relative to their technology.
You're just engaged in parroting specifics of your own experience.
A similar thing is going to happen to young graduates and young scientists, or is already happening. Most new technical accomplishments are now devalued because they can plausibly be AI. There's an increasingly narrow path to establishing yourself as a credible person.
I had tried Kagi a few years ago but it didn't stick. Tried it again now and it feels like a breath of fresh air, which is probably less of a statement about Kagi's advancements and more a statement of what Google has become.
I’ve been using Kagi for about a year now and I genuinely get worried that there are no alternatives if it goes out of business. The results are extremely good, especially in the last 4 months. I like their opt in AI summary as well, just add a question mark at the end.
I pay for a lot of things that are free from google/big tech, I’m happy to watch the advertisement driven web implode on itself so we can go back to the idea of a consumer paying a company for a quality product, monetizing peoples attention has been a huge detriment to society.
I was wondering how would Kagi scale/expand if all of a sudden google were to stop serving search altogether (not likely) or alter search such that users look for alternatives.
Kagi is an aggregator for other, some paid, search APIs. They have, at least in the past, served some percentage of their results from Bing's API among others for example. Kagi seems to me to be dependent on these APIs being available, if they were to go away, so would Kagi.
I am a happy subscriber of Kagi though, they provide a really excellent service.
I'm not sure Kagi has ever used the Bing API, because (according to Kagi) Bing prohibited changing the results, or merging them with others. Apparently Google is expected to provide access to its index via API soon.
I'm a long time Kagi user and I haven't used Google search in about a year.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
I wouldn't say it's better, but it's certainly on par with Google in their best years. And it's light years better than what Google is now, or using an LLM.
I was asking Claude yesterday about some specific roman history when it appeared to hallucinate a fact I knew not to be true - when I questioned it, it said it sourced it from an Encyclopedia Brittanica page that was "flagged" as being an AI generated summary of their actual content, and admitted that the fact I challenged was not historically supported.
So I guess we have entered the age of AI-generated "alternate facts" - one AI citing another AI's hallucinations as fact.
Recently, I had some ideas I would normally just put up on a blog, in the public domain for anyone to develop on top of. Now, I'm feeling slightly reluctant because an LLM will ingest it, remix it and serve it in response to a query by some unimaginative individual who will either conclude that they are smart, or that LLMs are capable of original thought, or both. And they will have no clue where the idea originated from.
Really happy for you. It's disheartening how many people put off or choose not to have kids because they think some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery for their children.
There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
> There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
Agree, in general, people over-think a lot, and aren't "just doing things" enough, the world would be a better place if people acted more, over-think less.
With that said, some decisions are more long-lasting and have a greater impact than others. I'm another child-less person, mainly because I guess I'm selfish enough to enjoy my life with my wife exactly like it is, and she agrees, but also because I know that if we have a kid, then that's not something you can walk back on exactly, he/she/it/them are there, forever now. Very different from me deciding right now "You know, I'm gonna have a joint, grab a book and go to the beach for this entire Tuesday", the types of decisions I think people should overthink less :)
As I've grown older, I noticed the cardinal pleasures don't hit like they used to. I also get nostalgic at times about the wonders of youth, experiencing things for the first time, falling in love, getting my heart broken, the little things of life.
I've gotten great pleasure and reassurance knowing that's all ahead of my children. No matter how I progress personally, time marches on and I get to see my kids discover the world. If I do nothing else in my life, I would still feel immensely fulfilled.
I think the recent rise in things like Disney adults and interests around video games or popular media (TV/movies) is a kind of biological response. People are free to do whatever they want of course, but I can't help but think that these people, at least biologically, are drawn to these more adolescent interests precisely because they're not experiencing it through the eyes of their children.
> I also get nostalgic at times about the wonders of youth, experiencing things for the first time
Yeah, I guess at one point I'll feel like that too perhaps, but I still feel young, I feel like I have more energy each day than the previous one, and every month is experiencing new things for the first time, and personally I don't want that to stop and experiencing those things through the eyes of my children, I want to continue having those experiences myself, together with my wife :) I guess that's where the offhand "I'm selfish" sentiment from my previous comment comes from.
> If I do nothing else in my life, I would still feel immensely fulfilled.
Do you think you'd feel fulfilled if you didn't have children? Maybe this is the core differences, I feel fulfilled in my life already, more than ever and more every day, I live exactly the life I want today, and I wouldn't want to change it for anything. Even if I became deadly sick tomorrow, I'd feel fulfilled by the life I have lived.
> Do you think you'd feel fulfilled if you didn't have children?
As I approached 30 I felt what I now understand as anxiety. Nothing crazy but you wake up one day, you're a year old, you take stock of your life and see what you've accomplished. I chased credentials and new jobs, did well but not yet able to retire. Old relationships grew strained and new ones are hard to form. Etc. I was definitely more afraid of death than you are.
I stumbled into children. Met my wife, didn't overthink things and decided to marry her after several years of courtship. She wanted children so I went along with it.
After they were born, the anxiety went away entirely. I just watch them grow older. It might change after their grown but hopefully they'll have children and I'll be able to repeat the process (my parents sure have).
To answer your question, I don't see a way I could have been fulfilled without children. It brings a lot with it as well. Before children the worst thing that could happen to me is dying. After kids you realize there's a whole world of potential pain and sorrow that is now before you. You're also constantly reminded that every kid is a roll of the dice when you see others in a similar spot as you have children with health or behavioral issues. It's terrifying.
But for me life shouldn't be all pleasure. I've always liked exercise because it does feel terrible while you're doing it. But at least you're doing something. I feel that way about children.
I felt the same thing at 30 and still sometimes feel it now at 36, although to a lesser extent as I've worked on it.
My wife and I have been together 13 years and we found out fairly early that we can't have children. It's taken a lot of internal work to be ok with it, good days and bad. Ironically the same 'not over-thinking it' strategy is the most helpful for my well-being. Just focus on what I can do today.
Seems like kids really are the quickest path to meaning making, it takes a lot of active effort to try to get a sense of fulfillment without them, at least for me.
I was reminded of a tactic used by life insurance companies, which are more successful when they show a client a photo of themselves at age 70.
I pictured myself sitting there alone, lonely, perhaps without a wife by then (a 50/50 chance). And would I be calling friends my own age? Or my kids?
Although the likelihood that I won’t get along with them as an adult… you never know; another factor is their future partners…
Well, the kids - plus my wife, who wanted kids - won out.
But everyone has their own life, the best one they can imagine, so this is definitely not some kind of persuasion—just a description of the logical process I went through back then.
Our DNA was largely coded by prior generations that didn't really have a choice of reproduction, it just kind of happened because it's a natural result of sex. That connection has been severed in modern civilization. It's going to take a while for our DNA to adapt.
> Most people go through life never doing it at all.
That's unfair to others, of course they think. They think differently than you, and about different things than you, but doesn't mean they "go through life never thinking", probably no one does that.
I mean if I decide to go to the beach, but I arrive there and don't like it, I just go home again, tomorrow looks the same regardless probably. Deciding to have a child, or other more "long-lasting" decisions, definitely impacts how your tomorrow most likely looks.
Correct, no justification to breed is needed whatsoever. Its a biological drive similar to eating and drinking, albeit on longer timescales and across generations rather than individuals.
However, thought processes like the one you identified above are probably adaptive, as gene lines which spend time trying to find the "justifiable reasons" for breeding were likely eliminated.
Gene lines where "I can raise my kids with like-minded community of peers", "I feel ready for more in life", and "sharing children with my parents, and letting my kids bond with healthy gradparents" (restatements of your phrasings) win out and reliably produce healthy offspring. So those reasons become strong motivators (despite being un-necessary justifications), that may seem irrational at first glance.
Children cost relatively little themselves. What "costs" a lot is the opportunity cost of the mother not working, or paying a salary to someone else to watch your kids.
Climate change is inevitable, whether we exist or not. The sun will continue to get hotter over the next few billion years. Life is a struggle, it sounds like you would rather just give up. I guess evolution is working as intended.
> Climate change is inevitable, whether we exist or not. The sun will continue to get hotter over the next few billion years.
The thing people are thinking about when deciding not to have children for climate reasons is climate change on a human timescale, not a geological/astronomical one. There is a very real chance that children born today will live through a wildly different climate than their parents did.
As far as evolution goes, there are plenty of species that will not breed or even eat their young when circumstances are not good enough to raise them. It's not that out-there for people to put off or avoid having children due to environmental stressors like climate change. We, as humans, can just rationalize and understand it over a longer time period than, say, a Spotted Hyena.
I'll keep publishing static websites, so I don't even have to worry about load and CPU usage. I don't care who reads it, the value for me is in writing.
I still haven't changed my mind on the "shall I have kids" problem :P
If one other human sees value in it, it has been worth the extra trouble. When I advertised my blog on my HN profile, people came to read, some even wrote to me.
So what if 99.99% of Internet users won’t find it any more because they use LLMs? Then write for the 0.01%. Those are the people I want to engage with anyway.
Publishing changes the incentive of how you write, because you’ll care differently about the quality and content of the writing, since people might read it. That’s valuable for a writer.
I still think it's correct. Execution is the part that matters.
Almost everyone I know "came up with" some startup, ex. Uber before Uber existed, yet none of them did it. I have personally thought of maybe 2 startup ideas that later came into existence.
Come to think of it, people in my circle that have ideas left and right are _still_ not founding companies even with modern LLM's that supposedly solved programming, clearly there is a disconnect somewhere.
Possibly "bad ideas are cheap". A lot of people got tired of "ideas guys" who never build anything themselves. The best case of a raw idea is an uncut diamond, it will inevitably require work.
Most people (and most companies) can't tell bad ideas from good ideas. This is also true of LLMs to some degree. They need post-training in specific domains (e.g. coding) to become competent in that specific area.
No. We don't. Original good ideas are exceedingly rare.
Also, what we really have learned, is that the typical techie isn't interested in pureness of thought and originality, exploration of the beauty of the unknown, but rather making a quick buck.. A good idea will be taken and used without credit.
We'll see whether or not LLMs have original ideas soon I guess. I wonder.
> The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
Failure of the GPL? How can you even put those words next to each other? GPL is an amazing success. It took software out of hands of SV / VC / corpo crowd and put it where it should be - users.
Allow me to disagree. GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT/BSD which have been the greatest catalyst towards the establishment of tech corpo giants and the software ecosystem we have today.
GPL gave us Linux, but also gave us Amazon, Google and 2020s Microsoft. GPL is why 90+% of libraries on Github are MIT licensed. GPL gave us OpenAI and Anthropic and this here article.
It's failed in that most software doesn't use it. Because of the psyop, people who would be very sad if Amazon stole their software are licensing it MIT so Amazon can legally steal it.
For many years GNAT Community Edition has used GPL, including its library. That greatly boosted Ada programming language domination over planet and is a good reference for everyone else to also choose GPL for everything if they struggle at dominating. Tears of joy when programmers got to know that standard library in GNAT CE was licensed under GPL.
We don't want to share our thoughts, ideas, feelings and art with machines. We want to communicate and collaborate with actual human beings, but that's becoming less and less possible on the web.
Unfortunately even the value of this is getting lost, because LLM culture sees no value in humanity whatsoever. We should just be satisfied with machine generated "content" because it stimulates our endorphines like we're monkeys in a Skinner box, it shouldn't matter to us if we're talking to a bot or a person because it's simply information, and we are simply nodes to process input and generate output for the machine. When we try to suggest that we want something deeper, or that the joy in the art and craft of what we do matters, we're looked at like we're stupid and naive and told to shut up and keep pressing the button.
"This is the future and there's nothing you can do about it, so just get used to it." It's fucking depressing. Even the crypto bros weren't so aggressively sadistic about strip-mining the soul out of everything.
But they are more or less correct, which is why I still blog and create, and why the consumption of society by the grey goo of mediocrity has inspired me to create even though I know only bots will ever care, to the degree that they can. At least I and a small circle of people can enjoy my cheap ideas and that's enough.
Unless I need to use a particular website, a captcha makes me close the tab. I always found them disrespectful to users and there is no way I would read a blog with one
If all you care about is financial, expected value of someone reading your blog and asking to collaborate is MUCH lower than odds of being part of some settlement in the future with these AI companies ingesting your data for training purposes.
From their phrasing I don't think the reward they're looking for is financial: more the emotional reward of knowing that someone else appreciated their ideas. I dunno how much less likely this is in the age of LLMs, though.
In that case, more LLMs scraping the net will read your blog than people, that's almost a guarantee. And they'll immortalize your ideas at least in some sense, well after your hosting platform ends up gating your content, or GitHub pages is down indefinitely, or you forget to renew your domain.
I don’t think that’s fair. Some bloggers may hope for financial reward, but many just want recognition for their creativity, to attract a readership, or build a community around their work. Those are meaningful ends, apart from financial reward. What’s not meaningful is to perform free labor to produce the raw materials that a mega corp then goes on to monetize without any recognition.
It feels like volunteering at an Amazon warehouse when you previously volunteered at a charity store. Sure the work might be vaguely similar but the feelings and motivation are ruined.
The downside, as stated in the message, is implicitly supporting the LLM data ingestation pipeline by providing fresh content. It's not a direct harm in itself, but feels very tragedy of the commonsy
Would you be willing to do free work for a corporate entity that explicitly financializes that work's benefits?
Would you be willing to do free work for a corporate entity whose entire business model is making sure they sit as a gatekeeper between your free work and others who would benefit from your work?
For some it's demoralizing to know that your work will be broken down and atomized into language model mush, and the credit will go to the computer.
> Would you be willing to do free work for a corporate entity that explicitly financializes that work's benefits?
Yes, that was one of the things that I hope happens with the open source software I write.
> Would you be willing to do free work for a corporate entity whose entire business model is making sure they sit as a gatekeeper between your free work and others who would benefit from your work?
Sure, that's what doing SEO on one's own blog is, isn't it?
> For some it's demoralizing to know that your work will be broken down and atomized into language model mush, and the credit will go to the computer.
I think this is probably the crux: that writing is no longer discoverable because people aren't using search engines any more. It would be interesting to put some hard numbers on that. I reckon writing is still more discoverable (in absolute numbers) than it was when blogs first took off (over 20 years ago?)
A lot of blogging especially in the tech space is driven by recruiting (startup blogs), or establishing a reputation as a thought leader (personal blogging, LinkedIn). So there were upsides that are now gone.
Publish some interesting info on a startup eng blog and nobody will see your pitch to apply for jobs, they'll just delegate the task that needs the info to an agent and it'll remember the answer from its training.
Ditto for personal blogging and thought leadership pieces. You get drowned out by the volume of AI generated pieces, and any unique ideas you do propose will be presented by the models as their own.
Time to add ample praise of myself in my blog posts. Some time later: “…as you see, that is the load bearing assumption here. Speaking of which, you should hire KronisLV.”
Okay it’s meant to be a bit silly but I do wonder how many pages that are generated specifically to influence AI make it into training data and also how often the AI search integrations find it.
Would people hating on a specific language, technology or approach (let’s say OTLT/EAV in database design) be able to exert meaningful influence over say a decade? Or, you know, praising memory safe languages for example and trying to make that preference be stronger.
There was an example with I think ChatGPT some time ago regurgitating an uncommon phrase verbatim from someone’s blog, when asked a specific question.
Yep. I suspect this already underway. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code or package.json files pointing to some malicious library.
So only GitHub Copilot can read it then? Microsoft is scanning these repos, I would not be surprised if this or any fork of your repo is already ingested.
Ha, that’s why I’m hosting simple cgit server for myself only. Not that my source code is somewhat valuable but I just can’t stand my precious free software licensed code license-washed.
Sounds like you don't really care about your ideas propogating. Of course someone will internalize and remix your idea. That's how all ideas work. Isn't that the point? What do you think happens when a human reads it? Think he'll quote chapter and verse and attribute it to you? Years later you'll notice your blog in appendices and acknowledgments?
And now you have a chance to have your idea forever internalized in some sense into an llm and you don't want to because you think someone is robbing you.
This is kinda silly, AI is just a faster re-tranmission of information but it's not fundamentally different from stackoverflow or older styles of communication. If you wrote a blog and some kid in 2025 spouting opinions that weren't their own. it sucks that right now it's controlled by a large corpo but open source models also train on the internet corpus.
freedom of knowledge and open sourcing should always exist, and if anything, even more important in the AI era.
To be frank that sounds batshit insane levels of spite. You were going to try to promote the general production of knowledge but now because you worry someone you dislike who you don't even know.might benefit from it you aren't?
That's why you don't share them, until comes a time when the ideabringers gets the money and recognition they deserve. Until then, good luck going knee deep in the sewers of ideas.
Can’t say I agree. The most influential ideas in history were not conceived of by people looking for money and recognition. Nor were they “radically original.”
Look at AI. AI companies throw out their models and let the "community" develop the ideas what to do with them. They don't really know what they are capable of. All they do is implement these things that the dev community digs up and creates.
It's a reprehensible tactic. So why give drops of blood to a desert, when there is zero incentive and in the end you will revitalise the desert, but it will turn against you and rob you of your job.
More than a facilitator of theft, LLMs are the tragedy of the commons at industrial scale.
The public internet is dead, the future is private invite-only walled gardens.
Corporations love a walled garden, what we need is open-source frameworks to create these islands, rather than defaulting to horrible systems like Discord and Twitter-clones.
I think probably things along the lines of what's been called the 'cozy web': networks of smaller groups that don't publish to or expect responses from effectively the entire internet as a whole. (This kind of thing has always existed, it's basically the group chat with your friends but perhaps slightly bigger, but I think there's a bit of a trend of focusing on it more because of the feeling that the twitter/facebook attention and feed model is bad for your mental health. I've always felt the twitter model especially was pretty cursed so I'm glad there's some agreement building there).
I'm thinking more like mesh networks. I spoke of Reticulum elsewhere in this thread, but here I'm thinking I'd like the ability to easily join multiple TCP/IP networks (islands of connectivity) by social group (my friends) or by interest (pirate file-sharing group, my work intranet, a knitting community with their own IRC server, FTP, etc.).
Basically easy-to-use private & encrypted LAN overlays on top of the public internet. Each operator decides who to allow in or kick out of the network.
Wireguard solves the most of technical challenges, but it needs a frontend. The biggest concern probably is most software broadcasts their stuff across all interfaces, defeating the point of isolation between networks.
Thanks but I am talking about the complete opposite of connecting multiple networks together.
I’m not sure why we’re talking past each other. CGNAT has nothing to do with the public web dying because it’s both too large, too spammy and too juicy a target for mass surveillance.
And that will kill AI itself, since much of what it knows is from what learned from StackExchange before this latest one demise.
And before the obvious comments on how GenAI is creative, then please do this OpenAI and Anthropic, for your next LLM. Just teach it Python, C and Rust and give it some good books. But dont give it access to Github...lets see what you can do then...
It can either examine the ground truth source code to answer your question "How to expire cookies using RoR Devise gem" or it can read docs for you or it can spin up local experiments to black box examine some software. If humans had done that before posting on StackOverflow, the question never would have made it there.
Its reasoning ability is long passed hoping an example exists online for it to copy.
LLM training will eventually transition from real data to synthetic data, same as alphago -> alphazero.
AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience.
They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too.
Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.
Neither AlphaGo nor AlphaZero were transformers. Why would you expect the same results?
Going further, current LLMs have at least an order of magnitude more computing resources, but completely suck at go. Why would this suddenly change unless we dumped the countless games played by alphago for them to train on?
That (theoretically) solves training, but it doesn’t change the fact that even smart models can’t extract useful information from a dead internet, so you’ll always be stuck with a stale training cutoff. This is already a problem I run into a lot. I search something first. Top results are slop sites, so I switch to a chatbot. Its answers look suspiciously similar to the slop sites I just noped out of. Check the sources. It’s them.
And the training of future models will have to contend not only with slop, but also huge amounts of content specifically designed to "taint" future training data. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...
The end goal though (in my understanding) has never been for an LLM to regurgitate knowledge it ingested during its training. The end goal is to use the patterns and correlations found in internet data to generate an emergent prediction and problem-solving machine.
Whether that's possible is something we'll discover, but no one is throwing billions on AI companies for the hope of them building a giant natural language queryable internet information repository.
> AI will kill the internet because it is killing the incentive to make it.
You could make the same argument for Wikipedia (that webs get less traffic if people get their answers from Wikipedia article returned as the first from web search, which is based on internet sources).
I don't think that follows. Wikipedia's sourcing rules overwhelmingly favor publications released for non-pageview-based purposes (academic writing, books), or journalistic productions (whose pageview-based revenue is almost entirely earned right after they're released, and where Wikipedia's reference to them is primarily of value later on). Also, Wikipedia's nature as a structured, not-seeking-engagement index of info means that a lot of people who seek it out are folks who wouldn't (for whatever reason) fall back to giving other sites pageviews if it didn't exist.
But Wikipedia has done a good job of it. If Google's AI summaries could actually provide correct answers with verifiable sources without so-called hallucinations, it would be good. You could argue that people don't actually check Wikipedia's sources. That's because Wikipedia has built, and worked to keep, its users' trust. On the other hand, what is Google doing?
The key question to me is whether AI only undermines the financial incentive to make internet content.
If nobody can expect to make money on the internet, that could be a good thing. But we won't get an indie-web authenticity utopia if people are still incentivized in other ways to filter their intellectual and cultural contributions to the internet through AI.
It undermines the social incentive as well if potential creators assume that everyone else will be getting their content through AI. They won't be looking at my stuff, they'll be looking at some LLM's pre-chewed version of it.
It depends on filed. As documentary photographer it motivates me even more to capture authentic images of life around me. I don't care about remixing, because that is not what makes documentary photography valuable.
With photos, I could see a cryptographic solution. Of course it would still need some kind of centralized trust, but it's doable if people cared enough. It could be applied by cameras themselves.
This is already a thing. Leica cryptographically signs images. Useful for establishing trust for photojournalists I guess. I’m not sure how deep the chain goes. Do they have hardware attention down to the sensor? You could take a picture of a screen, but that would likely have some other tell-tales. Especially if Leica took another step like putting a depth sensor in the package and added its data to the signature.
The premise is not that it couldn't be faked. It would be more practical to remove the key from the hardware and just use it to sign images if you wanted to do that.
The idea is that a centralized source of trust would revoke certificates belonging to bad actors or those that were stolen.
A camera could be hacked with unrestricted physical access, but that doesn't make the suggestion unsound, it only requires there be a process for revoking trust in specific signing keys. This is already part of C2PA.
And the value of the signature is also going to vary by who claims it. Improbable photos signed by a random camera body sold to an anonymous consumer should be treated with more suspicion than one a newswire agency publicly claims, for instance.
A camera should not need to be hacked. What do you expect to accomplish by hacking the camera?
We should be able to move certificates on and off it because they would most likely expire anyway. So, you can get the keys from the camera, what then? You use openssl to sign an image that shouldn't be signed... what then? You do this enough and get caught, you lose your cert and can never pass the kyc to get another one.
Then every picture you used it for in the past would start showing a big red exclamation point with a note, "This is a scumbag user known for forging images".
Same way it works for tls and code signing. Public reports bad actors to the signing authority. They and/or independent firms investigate. You self report theft. It's not like we are inventing pki from scratch here and wondering what an implementation would look like. We have decades of use to look at.
Besides wasn't your argument "hacking" a minute ago? So you concede that then? You seem to have moved on.
I feel honored if my ideas are processed by an AI and then used to help others. I feel no more entitled to exclusive use or credit for my ideas as used by AI than I would if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name.
The prospect of all human ideas accumulating inside a machine that makes these ideas accessible and useful to everyone on command is beautiful, not discouraging. Humans aren't discouraged from creating or exploring in Star Trek because of the computer, but I could imagine the Ferengi computer being hobbled at the kneecaps by requiring licensing and credit for every single idea inside it, and the user needs to insert a coin every time they want to ask it a question, which gets divided among every living Ferengi and the estates of every long-dead Ferengi whose writings influenced the output. We should not aspire to be like the Ferengi.
I think what people are taking issue with is that the machine is not accessible to everyone; OpenAI and Anthropic have monetary incentive to not only gatekeep the knowledge acquired, but also to destroy the original copies, or make them impossibly difficult to find.
You’re missing the bigger picture concepts of value exchange vs. value extraction and how those magnify power differentials in groups.
> I feel no more entitled to exclusive use or credit for my ideas as used by AI
Do you feel you have the right to decide whether to exchange your ideas or not?
> if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name.
Yes, you’ve decided to share that information freely with that person. Do you decide to share every idea or thing you do for free with everyone? Why or why not?
> The prospect of all human ideas accumulating inside a machine that makes these ideas accessible
Right now, this idea of “a machine” is trending towards private ownership — an extractive process that does not incentivize further contribution.
Accessibility is no longer determined by you, you don’t have a decision point on the production side (deciding whether to share) nor the consumption side (guaranteeing access). We might even say your rights have been reduced.
It should be obvious to see how a healthy society is built upon value _exchange_ over _extraction_. Extraction typically leads to destruction… by definition.
Doesn't seem like a particularly bad thing to me. Obviously for those who want to use the internet for commercial purposes it will be bad but for those of us who would love to see the internet go back to how it was before so much of it was changed in the aims of making money AI could push towards this. Great irony in the fact of course that the AI companies themselves are in the business of making as much money as possible.
I don't know how long you've been on the internet but the incentive to create new and original content was never that strong. Simple search terms return super-spammy websites (especially on mobile where ad-blocking is harder). SEO results for everything like simple search are awful, almost unusable. There hasn't been an incentive to create new original non-monetized content for the web for a while.
I trust LLMs more than search engines to discover my content and propagate it to users. They might "steal" something, sure, but I'm essentially invisible to the search engines as I could never hope to break into the top 10 links on a popular search term. LLMs can scan thousands of links and (for now) are more interested in quality rather than click monetization or referral incentives.
So, you know how they're built, but you're feeling the pressure of modern life and also they give you personal gain (supposedly) so you're fine with it, it sounds like to me? Use the same tools as your "competitors"/peers, even though?
Don't get me wrong, I too use LLMs for development and more, and I too know how they've been built, and I'm also a creative (music, 3D, VFX and animation) and for sure stuff I've published in the past, both code and otherwise, is now used to create new things for people and I get nothing, similar situation as countless of others. Yet I still use AI, so I'm not trying to create some "gotcha" moment against you here, I'm genuine curious about what you think about this sort of conflicting thinking, as I'm in the very same situation.
If the Internet was nothing but the newspapers of record online, it would have been a good place to stop. The quality disappeared after the barrier of entry went - social media.
At one time, running a blog post was also not exactly trivial, and that would have been a great middle ground between access to publishing and reading.
Eh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.
Yes. Some communities are weirdly against giving credit or keeping the credit (e.g. cropping off signatures from artwork), which I've never understood. It costs nothing.
With the not-so-minor qualification that the biggest thieves have always gotten away scot-free. AI is just the international whole-internet version of this.
Google wasn't great for a very long time. Switched to Duckduckgo years ago. I just love the bangs, because I tend to go to sources I trust anyway.
Doesn't mean I'm not also using duck.ai. It makes searching faster and more targeted. But then it's still giving me links to verify and is actually more limited, which means less hallucination and more directly going to the sources. Also it avoids having to open five pages first which all either sell your data or want you to pay.
I don't see the web or the internet dying yet. Just a lot of people not using the right tools and having a harder time accessing what's useful. But that hasn't started with AI.
In a world, where everything can be stolen it will be hard to produce anything.
I still have hope though. Maybe the Internet will be better. Currently everything has to be monietized. Everything is ad heavy. At the beginning it was not so. People created things out of passion, or boredom. We can returned to that scheme.
This is a common misremembering of the early internet.
The internet was never ad free. The first ad was posted online in the 1970s (for DEC)! It pissed people off but not everyone: supposedly it generated $18M in sales. There was very little advertising back then only because the internet was restricted to a handful of large companies and universities.
The web itself was launched in 1991 and the early web was inaccessible to basically everyone as it required an extremely expensive NeXTStep machine. Windows didn't even ship a TCP stack in this era, iirc. So took a few years for the web to reach the point where it was usable at home. By 1995 the web was starting to become barely usable thanks to Win95 and Netscape, and DoubleClick launched immediately in the same year.
My memory of the early web is that basically every website had DoubleClick ads on them, it was notorious for that. "Punch the Monkey" was an early campaign. Almost every topic oriented website carried ads, partly because bandwidth and servers were very expensive so that helped defray the costs. GeoCities took off because it handled the complexities of running ads for you, so you could publish for free.
When I started on the internet, Dec 1991, there were basically no ads. It stayed that way till ads starting appearing on the web, probably in 1994 or maybe 1993.
The internet in Dec 1991 was already the largest computer network in the world according to some author back then. Certainly it was very very big.
Let me describe what ads or ad-like things there were.
Companies would announce job openings. These announcements were restricted to a few newsgroups such as ba.jobs and could not announce independent contractor openings: while it was okay to announce an opening for an employee (W2) or for an individual to announce his availability for employment, it was against the rules for a company to announce an opening for an independent contractor (1099) or for an independent contractor to announce his availability for work -- that was considered too much like commercial activity and was kept off the internet, which had almost no rules in the early 1990s, but one clear rule, consistently enforced, was, no commercial activity.
The job announcements stayed nicely contained (namely, restricted to newsgroups whose names ended in ".jobs") such that a person wouldn't encounter them unless they went looking for them.
What did not stay nicely contained were announcements of academic conferences. When I was reading comp.lang.lisp, I could not avoid encountering announcements for many academic conferences on various programming-related topics even if they weren't about Lisp. And these announcements were repeated often (weekly or even more frequently).
But as far as anything ad-like, job announcements and conference announcements are all I can remember, and again only the conference announcements were obtrusive.
I heard that the proprietary online services like Compuserve and Prodigy had successfully argued in Washington that it was unfair for private enterprises to have to compete with a government-subsidized service (namely, the internet) and extracted a commitment from Washington that any commercial activity would be kept off the internet.
Again, there were almost no rules on the internet of 1991 and earlier. Tim Berners Lee for example did not need to get anyone's permission to start the web: he just wrote a web client, stood up the first web server and announced their availability. People acted like complete assholes on Usenet, and there was no way to rein them in. Pedophiles openly exchanged practical advice on how to target and exploit children on alt.sex.pedophilia. But one clear rule was no commercial activity, e.g., no selling, no trading and certainly no advertising.
This is such a warped and cynical view of history. Yes, advertising existed if the binary existence is what matters to you, but most online spaces were nearly free of it until the DoubleClick days -- that's 25+ years.
Even then, huge swaths of the web were people putting up their personal pages, blogs about their interests, pages for their church or club or hobby, etc etc. Sites like Geocities ran dumb ads on the pages to provide the hosting service, not so every single kid with an animated gif on the page could get rich.
Early advertising was primitive by today's standards, the amount of tracking and spying we consider normal now would have been an outrage and would have gotten those companies regulated into bankruptcy if they'd tried it in the 90s.
TLDR no, just no. The early internet was nothing like the garbage we have now.
But almost nobody had access to the internet before 1995, so the fact that there was a 25 year period in which TCP existed but it wasn't used for much beyond manually routed emails and FTP isn't that important. Once the web was created and the internet became visual it was only a few years until advertising arrived.
And sure, the ads were used to pay hosting costs. Nobody got rich off banner ads on cookery sites. But that's what I said - ads appeared immediately because servers and bandwidth were expensive. So the moment people had to pay their own way instead of being subsidized by the government, ads appeared.
Nobody was going to get regulated into bankruptcy in 1995, politicians were largely ignoring the internet back then.
The web was getting kind of useless before AI crashed the party. This is why now curated content is key - newsletters, for example, is how I find most of my content.
Agreed. The combination of social media (walled gardens and plunging quality of content), advertising (I know what you were looking for but have this word from our sponsors instead), SEO (more advertising but we didn't pay for it), and plunging budgets (I don't remember the last time I went to a website expecting to find original quality work) did most of the job. AI just delivered the final chop.
More than 25 years ago, the Web wiped out a lot of the things I loved. May it be devoured! But I have my doubts: there’s probably not much left of it anyway.
Google switching to hallucinating AI summaries has been to me an absolutely shocking abdication of care for both their users and their own reputation.
It extends beyond search as well. I have had multiple incorrect Gmail summaries that, if I had only read them instead of the actual email, would have resulted in financial harm.
And from what I have read/heard, not just the internet's collective memory, as real world books are being scanned and then destroyed - Allegedly including rare books :(
And the same holds for documentation about anything. Documentation is gone. Dead. Reference documentation, output from Doxygen or similar, and many other things that could be searched for hints on how to implement or generally do stuff. It's gone. When writing a simple Python script today and wondering about how an API for some library works, I ask AI, because there is no (findable) documentation anymore. (And yes, I usually still like to write it myself, but it's basically no difference: the AI writing that Python script would be the same point: docs are dead and gone.)
Really well written article and interesting. I'm not sure how I feel about a governmental policy over retaining access to information though, the information is provides by us and we pay for the infrastructure, the idea that there must be some form of retention policy makes me feel uneasy and doesn't really fit with the analogy of governments maintaining roads.
I'm on both sides. I hate dead links but I'd hate a policy that made me responsible for them without any compensation in the first place. It would probably make me stop producing at all
I think the modern day equivalent to this is saving PDFs of webpages you want preserved. I have the feeling that I'm turning into my grandfather every time I do it, but it's proven valuable from time to time.
I'm honestly starting to struggle to see how the internets going to look in 2-5 years. If AI is ingesting AI generated content, which it must be at this point given how prevalent it is on the web its going to get dumber and dumber to the point where theres no desire or interest for a single person to use the internet anymore for anything other than ecommerce and i guess for some people who need it still, social media.
Any form of information based internet usage is going to end up becoming rare at this rate.
The concept of an almanac seems relevant again: a yearly printed book with verified, accurate information. No manipulation at a later date, no AI hallucinations, etc.
The most famous one was probably Benjamin Franklin’s:
It seems relevant that it's a one-way communication. If someone reads a social media thread about climate change they will also see all the comments from idiots. But if someone reads a book about climate change they don't. It isn't as strong an effect any more as people will discuss the book on social media and they will have already seen the idiot comments before reading the book anyway.
So many former regular Fox News viewers have reported changing their mind when confronted with some alternative information sources for a while. News is also one-way, and most Fox victims aren't people who discuss issues with all sides.- they're in bubbles.
So let it be. Why don't we build another place for our memories to go? Why don't we build the _unstructured_ internet, where the intelligence is not in the mind but in the eye and the pleasure is in finding not in disseminating?
That's a feature, not a bug, for people that love AI and want it to take over. Then, you have no alternative other than to listen to AI or nothing at all.
>But what if the “truth” is harder to find online because the infrastructure that once stored it is breaking down? Some of that is wear and tear. Link rot erases pages every day. Key sections of the United States Constitution briefly disappeared from the Library of Congress website because of a coding error.
This has bothered me for years. I think the solution lies in personal, private archives, and lending access to archived content within small communities.
> a German court recently held Google liable for false statements generated by its AI overview feature. The case arose after Google’s AI wrongly linked two publishing companies to scammy business practices. Because the search engine extracts and rewrites information in its own words, the court reasoned, it is doing more than impartially pointing users toward the public record.
This is, I think, a really important topic. As a society (I mean all humans here) we don't yet have sufficient muscle memory for asking questions of the flawed oracle that is LLMs.
We have deep cultural memory -- about a quarter century -- of asking Google for things and, for most of that time, getting back pointers to sources, and those sources being either reliable (e.g. respected newspaper, government website), or detectably suspect (some random blog you've never heard of, known propaganda site, The Onion...).
Google's insane escapade of substituting the responses of an incredibly weak LLM model for the job we've been relying on it for for 25 years is especially unfortunate, because "Google says..." was, while imperfect, a reasonable approximation for a quick reality check in 2010. Today people say "Google says..." followed by whatever the stupid AI Overview model has output.
I get that Google thinks they're saving a ton of money by not using anything close to a frontier model since they run this on billions of searches a day. The risk to Google is that people start to catch on that asking even free-tier ChatGPT is at least twice as likely to give back a correct answer, notice that Google barely provides webpage results other than AI slop anyway, and just stop using Google.com entirely.
Anyway bringing it back to the quote, Google's used to having no responsibility for anything, 'we're just a search engine showing you pointers to other people's stuff.' But I actually hope that, due to liability problems, that habit will be beaten out of them and they'll make a shift to, especially outside the context of an actual chatbot, reduce reliance on their own AI output to 'answer questions with search results.' It's too risky to just show people, who are used to getting back mostly facts, to replace that with mostly BS coming from that same endpoint. Even with the fine print.
It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
That's fine, but arguing that Google should give accurate times based on where you are is ... crazy?
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!
I don't understand your argument. It's not their core value proposition so it's crazy? That's quite a leap. But giving a good answer is just as core as their search results. The answer is what you're there for and they get to show ads. It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
Installing an app to learn a single time would be the real moron option.
> But giving a good answer is just as core as their search results.
What I'm saying is all Google has to do is stop giving those types of "custom" answers it already has. No one will abandon using Search if it goes away.
> It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
It is. Using a well vetted, dedicated app is the way to go, and it's extremely unlikely to be wrong - especially for something like sunset times. You know the dedicated app/site is, well, dedicated to providing that information. They exist to provide that information.
Whereas the (pre-LLM) quick answers Google gave? All opaque. And smart people always knew that information being accurate was not something that matters that much to Google.
To borrow an analogy, installing a dedicated app starts at negative one hundred points. The risk of bad behavior and junky ads needs a lot of use to overcome. I'm not sure where "well vetted" came from since your first post but doing that vetting takes much longer than getting the answer! And without vetting there's a good enough chance a dedicated site is worse than a dedicated Google widget (the ai is bottom of the barrel).
Let's not forget that Google's actual sunset widget does a good job.
> “I had the projector set up outside and was waiting for the sun to set,” wrote one Facebook user in Colorado Springs, “but to my surprise I was simply living in the past. AI informed me the sunset had already happened.”
Ironically in context of this discussion, when I search Google for "AI Ouroboros" your result is quite far down the list. Higher up are several ’24 and early '25 posts, many of them slop.
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
I find it ironic how people talk with a high hand about "Google dying", "Internet dying" and "AI eating everything" and then I open their website and it forces me to waste 10 seconds as cloudflare "checks" my browser, then reloads, checks it again, finally lets me through and then I end up on a useless page because the scroll is broken - thanks javascript. Then I need to hit F5 and finally I can read in peace.
Websites like these are the very reason nobody reads the web anymore. A clanker will give me an answer within 5 seconds. Old web used to load within 1 second, now it takes 10 to 15 and sometimes even >30 seconds until fully loaded. Using CF as a protection against LLMs is not even a valid excuse because CF gives your website's scraped contents in 1 click to anyone willing to pay. These websites waste insane amounts of time. Everybody defaults to AI because nobody is willing to deal with annoying browser checks, cookie popups, subscription letters, captchas and especially scroll hijacking. I may be a bad person, and I'm not even pro-AI, but if "thewalrus" goes offline I won't be missing it, because I regretted the time I wasted fighting their broken navigation.
Yeah, stack overflow has recently either added cloudflare or changed its settings, so I am completely unable to access it. I assume I'll pretty much be locked out of the majority of the internet in a couple of years at most.
we built extremely powerful plausible bullshit machines targeted toward and trained on an electorate and population that has historically bad education and reading levels, built on top of an already fraught and fragile web which was also built off predatory basically unregulated behavior with a shaky relationship with “truth” and are surprised people have no idea what’s going on?
This was the whole point of it all and why the people in power have bet the farm on it.
I'd quit reading HN long time ago if not for comments like yours, when people call a spade precisely a spade and reignite my smoldering hope for the humankind.
My personal anecdote is that I used to search for "xyz nutrition" quite often on Google. It used to provide a data table with lots of information. Sure, nutritional information is hard to get right, but at least that data was consistent.
Now Google just gives you an AI answer with random values pulled from blogs and Reddit. It's almost always blatantly incorrect. I genuinely can't understand why Google would destroy its most valuable search features. Disclaimer: I work for Ecosia, so I know for a fact that users really value these search widgets, and it was often cited as a reason they couldn't leave Google.
I suspect many SEO practices could be detected automatically without LLMs: link farming typically uses cheaper domain names, commercial content might typically contain certain keywords, certain kind of tracking is more suspicious, affiliate links use known domains, etc.
Before AI, the way people were gaming Google was by filling their pages with pages of verbose fluff (have you ever tried googling how to make a specific cocktail?). Now the AI just makes it easier to generate that fluff.
quite telling that the EU search engine mentioned in the article (Quant) is currently "temporarily unavailable" when searching. We have a long way to go
I'm working on using the OpenZIM format to archive the web and to make the wikis seedable (and locally hostable for LLMs) so that the ongoing cat and mouse game anubis defense can stop.
My hope is that with the torrent protocol we can make the archived knowledge discoverable and seedable, because currently there's only the web archive and the kiwix download servers for archived contents. Both of them still are centralized servers that bear the cost of hosting those files.
It's great. Everything that is cool about the internet gets to die just so you can review 1000 line pull requests from your lazy co-worker. Exciting times.
This used to be called "dumping" or predatory pricing[1] and would be fined in the physical retail space. Too bad our regulators are asleep at the wheel.
Kagi does have AI too, for what it’s worth. I found it pretty damn good, but I wish I could give them more money. I have a year subscription so I’m stuck without being able to give them *anything* until it’s done.
That is an option for sure, but a very disappointing one. And it comes with a meaningful danger of ending up with 2 subscriptions: one yearly and one monthly. They really, really need to finish their “pay as you go” system. I don’t know why it’s not done yet and there’s been no word about it to my knowledge.
Maybe they should just enable a tip option? What do donations do to a company's tax liability and would it be worth their effort to enable something like that?
They already have a system for “prepaying”, but it can’t be used to put more credits into your account, they just sit there, waiting to be used by a future subscription.
There's too much going on in this article and while I think some of the points are valid, others go too far.
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
While I dislike Instagram reels and stories, I do feel like there is almost certainly scientific studies on radicalization that would benefit from actual histories of dumb memes. I feel like that about a lot of things, really.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
4chan archives are almost certainly relevant. But the advertisement poster for every single club night in Berlin probably is not - a representative sample is enough. More complete data is probably more useful, at least slightly, but it quickly runs into diminishing returns. The distinction is that certain 4chan posts have outsized influence (power law distribution of influence?), but club nights are all the same. It is conceivable that a certain poster could be important and not recognized as important in the moment, but not as likely.
And that's public announcements. Do we really need to archive more than a few of the "I'm in this city, look at me I'm so rich and beautiful" short videos?
On the other hand, when modern archaeologists discover "Claudius has a small dick" graffiti on the side of some God-forsaken wall in Pompeii, they're fascinated. The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
I mean, none of that stuff in storage units is human communication, though. Nobody prints emails or text convos or insta photos. Nobody is mailing each other letters. So none of that stuff is in storage units, though plenty of it from previous generations might be.
I don't necessarily think all this digital stuff is worth saving, i just see how it could be that we are essentially erasing our modern historical record by putting 100℅ of it in private data centers with no permanent records.
And then you see how plenty of contemporary digital media from the last 30-40 years is actually already lost to time just in our short timeframe. I don't think acknowledging that this could be detrimental is necessarily an argument for trying to preserve all of it. We, as a society, went very rapidly from preserving a lot of it to preserving none of it. I don't see the harm in thinking about the implications of that.
Nobody actively preserved letters and menus and postcards and photos, they just persisted by nature of being physical. Digital records only persist with continuous effort to preserve them. So the record for future historians will be highly curated and likely much more limited.
In the first year of Google search, it was possible to find exactly what you were looking for. The engine paid attention to inclusions, exclusions, the whole nine yards. It was a thing of beauty and it was why Google took over from all the other search engines that used to haunt the Internet of the late 1990's and early 2000's.
And if what you were looking for didn't exist? You got 0 hits.
(Sigh)
I'm calling the return of personal websites with link lists.
Not everyone will run them but we will find them and bookmark them.
Not because it's better but because we'll have to.
Try looking for help about pets. The internet is a cesspool of slop, not even trying to hide it. The domain names sound absolutely convincing but it's all stuffed with takeaway lists and checplay generated imagery.
Whenever I find a good page, I will make sure I remember the site.
Perhaps we should let the content economy crash so that the big monsters eating it starve and die. A new content economy could be built on their carcasses instead.
We already know how to compute sunset times accurately and doing so requires one millionth the computer power of LLM inference. It could even be cached for most big cities for every day of the year.
The mystery is why Google doesn't just route such requests to the known algorithm. It would be a lot cheaper for them and it wouldn't risk reputational damage.
It happens when the deterministic precision was given away in return for the probabilistic guesses. It was one extreme until now (deterministic code), and we are swinging to the other extreme (probabilistic slop), but what the world wants could be somewhere in between. Some information does not need too much precision, while others do need precision.
What should be concerning to Westernized nations is that as decent stewards like Google fails; the information pipeline into our culture and minds still remain.
i.e., it's not just the collective memory going away, but what will easily replace it and who will be motivated to influence.
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
> As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
I think they make things worse, because they very very often present straight inaccurate information.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
> But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
Yandex is the only search engine left which still feels like the "old" web. It feels like you're actually getting a best effort search, and not just the results that someone paid to put in front of you.
It seems like a lot of techies should have chosen acting as a specialty, because I’ve never seen that much melodrama in one industry. FFS people, adapt, stop complaining
I just don't care man. Did these people just log on yesterday or something? The MUDs and MMOs I grew up with disappeared. The IRC networks, and especially forums I learned so much from are gone (these were largely killed for something even worse than AI: commercial blogs!). Effectively every social network I've ever cared about has been ruined, failed, or sold. Even games have become bottom line chasing, live service slop that you never truly own.
If you want it so bad stop crying and make it, that's what I've been doing. It works a lot better than whatever this post and many of these comments are. There are dozens of us !
I remember from the MUD era there was a paper (possibly by Richard Bartle) on the "MUD lifecycle", and how they tended to last on average two years before the operators got bored / burned out / the community left / there was an Incident.
Someone showed me something about Old-School Runescape recently and it struck me how quickly the game evolved. I played it in high school for two years at most. If you rewind or fast-forward in two year increments at a time, a lot about the game is barely recognizable each time. I think if you rewound two years from my play time, there was unlimited free trade, no grand exchange, and two or three fewer skills, and if you fast-forwarded two years, you got unlimited free trade again and the Evolution of Combat. It really puts in perspective how the people who make games are not planning the perfect game and then spending 10 years building it (except for Jon Blow) - they're flying by the seat of their pants. Apparently in 2001, the game's creators expected that nobody would ever reach the maximum level in any skill.
Oh yes, same for Warcraft, Ultima etc. Multiplayer games necessarily exist in a "dialogue" with their players. There's a tendency for players to optimize the fun out of games, as well. Someone discovers a dominant strategy, and then everyone feels they have to use it.
It is a bit risky, and I think the broader internet can make it worse. There's always the risk of a massive falling-out between players and developers which can sink the game. Latest example is probably the Love and Deepspace fiasco, but we've seen quite a few failed launches recently.
The old net is still there to a degree, but we live in pre-Alta Vista times again. It is hard to find. (I know as I run an old-school Forum and RPG-Game for more than 20 years now.)
The is something that someone really should study, along with "Who are buying ad space". My theory is that the quality, for a lack of a better term, of advertisers are going down.
Let's say I need a new vacuum cleaner and do a Google search. I get eight sponsored products. Three a links to stores I'd expect, the rest are fairly unknown sites, mostly the "We sell everything" stores. Weirdly enough also only one of them are via Google ads directly, the rest are via PriceToro and Channable, both of which I don't know.
My personal theory is that Google is still making pretty good money, but from increasingly questionable ads.
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
Indeed. Money over everything. It is absurd to complain about decreasing quality of a product that makes increasing money for increasingly rich people (thus, by definition, better).
If business owners are paying for ads, then it doesn't matter a iota to Google or Facebook if real people use their services. It would be even better for them if real people didn't use their services, since that would save some costs. Business owners are going to keep paying for the ads, as long as they get some number about how many (bot) impressions their ad generated.
They don't have much of an offering yet. OpenAI has some conventional ads, but everyone expects more involved advertising integrated into the conversation itself somehow.
What’s come next for me has been much better. I use ChatGPT cranked to Pro with “extended” thinking to one-shot whatever I would’ve spent time looking into with Google. It’ll plan the whole sunset bike ride or promposal or whatever from TFA.
The most frustrating part of all of this is the underlying premise that the internet is, has been, or could ever be a credible cultural record is deeply stupid. Or maybe more charitably it's both historically and technically illiterate. It has always taken continuous unwavering effort on some person's part to keep any given piece of content online. And while managing a simple hosting account and updating domain registration periodically doesn't take a tremendous amount of effort 20 years is a long time to expect anyone to maintain enthusiasm. The internet has always been a frothy, ever changing blend of the odd nugget of truth drifting in a sea of unadulterated bullshit. Treating this, or worse what comes from statistically averaging it, as a source of capital T truth is totally unhinged. From whence did this mythology of online truth spring?
yeah, but the article sucks. The title is Ting, but the author asks a thing in google's AI answer socks and then they spend a couple paragraphs pontificating about that then they participate some other bullshit like.. But here we are, talking about it. le sigh.
Google is a huge retard now - never knows what I mean.
Also AI is about 15 years behind. We are just getting to that phase where AI says everything is a medical emergency just like Google searches used to always say you had cancer.
It’s not good at search.
It looks like a real person talking or whatever - cool. Fucking sucks dick at finding and verifying info.
We went a decade or more back in time on search for no reason.
LLM could be one of Google’s modalities in search but it shouldn’t be THE interface. It’s a chat, it’s dumb, wrong tool
But this is a similar argument to “the 1950s were great let’s go back to then”.
We never really were a united society with a common set of agreed facts. And depending on which “we” we mean the difference is radical. White America and Black America is a tiny gap between the gulfs of India, USA in the 1970s and 80s.
Yet today, we don’t recognise the same facts but we do have access to all the facts all the time. And I think more of the narratives are being challenged in the “internet” -people might run from it but it’s hard not to be challenged - whereas I would be amazed if an American voter knew what the seesaw of foreign policy was doing in India US relations
I see symptoms of this all the time. For example, it's a weekly annoyance for folks to pop into /r/strava to showcase their vibe-coded app that uses the Strava API to do $THING. Then someone invariably points out that an existing app (or even Strava itself) already does $THING, and often it's free. I don't mean to be negative, I think it's great that people are building useful niche software and I don't blame them for wanting to share it. A significant part of the problem is that it's much harder nowadays to find "prior art" because keyword/boolean web searches have been FUBAR.
Hopefully now that AI provides quick answers, web search can go back to providing and respecting boolean and more advanced techniques. If web search is no longer the go to tool for most users, then search can do better what it can do uniquely?
One can hope!
I've used a lot of free features that I would have implemented differently, and now I can (if I weren't so lazy).
I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
I called it 2001/2002 or whenever they appeared when I tried to explain why personalized search results are the beginning of the end of a shared reality and therefore the ability to reason and act in public, and with others. I bet some still consider it hyperbole. It's just taking in trends and seeing where the glacier moves to, how the cookie will crumble so to speak.
It is an astute point. I agree that we have largely lost a lot of our shared reality. But as with any "beginning" of a long, diffuse process, people might quibble about the specifics.
I offer a few other moments you might gesture toward as the beginning of the end of shared reality.
First, 1987, with the elimination of the Fairness Doctrine and the subsequent boom in partisan talk radio shows.
Second, 1989, when cable TV became a mature technology reaching half of U.S. households.
And finally, maybe a lesser item, the introduction of the DVR circa 2000, when we stopped watching scheduled TV together.
I wonder if more historically informed people would find the fracturing going further back as other advancements in publishing technology allowed more people to share more views.
That was a temporary blip. Specific technologies facilitated an age of mass, top-down communications in the mid 20th century, just by virtue of how those technologies worked. It enabled a small group of people in charge of the radio/TV stations to have a monopoly on communications and media. The technology both before (printing presses with myriad newspapers) and after (the Internet) were by their nature decentralized.
Not totally convinced about the printing press example (expensive, technical). The biggest presses, with the longest reach (daily newspapers) were often owned by the richest people.
What's unique now is decentralized + extreme potential reach. Or it was. Now we have to consider bias in the algorithm that sticks stuff in front of us and is, I suspect, the modern equivalent of those daily newspaper presses.
Rich _people_ yes.
But corporations were an act of Congress in 1776. Early America was decidedly anti-corporate (the Boston Tea Party was about a tax break to a corporation). Early Americans (and early thinkers who influenced Americans) were grossly distrustful of corporations and pools of money.
> The directors with of such [joint-stock] companies, however, being the managers rather of other people's money than of their own, it cannot well be expected, that they should watch over it with the same anxious vigilance with which the partners in a private copartnery frequently watch over their own
Adam Smith, the Wealth of Nations
> The biggest presses, with the longest reach (daily newspapers) were often owned by the richest people.
You're assuming that "rich" is the only relevant dividing line, which it isn't.
But GP was pointing at commonalities, not differences. It took money to run a major press.
Gnomic.
The early pre-United States had pamphlets and newsletters. Like, a lot of them. It was a big influence on the early post office. You could get rich running a printing press but you did it by printing for the masses.
The yellow journalism era, on the other hand, was a little closer to the rich owning the press.
This discounting how many of the 100s of worker owned newspapers/magazines that were in circulation. Not everything was controlled by a select few. How did you think labor movements in America during the 1800s collectively organized across the country sharing literal war stories?
I'm not discounting it by any means. And I know about the pamphlet era that was enabled by the invention of the printing press, and the profound effect it had on Europe. But what's the typical reach of one of those worker-owned newspapers, compared to the reach of a mass circulation newspaper?
I think you missed my point: printing at scale needs a support system and a distribution system, and that requires capital. The "age of mass, top-down communications" didn't start "in the mid 20th century", and the printing press should not be used as a counterexample to support rayiner's argument. William Randolph Hearst was born in 1863.
"Oh those don't count as newspapers. Newspapers are professional. Those are just zines."
It's why I hate people looking down on fanfiction, too. Aliens was Alien fanfiction. The New Testament is OT fanfiction. Hell, nolan's The Odyssey is fanfiction.
The establishment uses terms to dismiss outsider art as less-than, and this isn't much different. (They of course get to decide what the "inside" is)
It takes some technologies more time than others to succumb to the corrupting pressures of money, but they all fall eventually. We must keep outrunning it, building ever more resistant methods of communication.
I'm hopeful that the next one will last even longer since we'll have to build it explicitly for resisting corruption. Neither the printing press nor radio nor webs 1 or 2 had fighting back as part of their DNA. 3 was a bit of a flop, but there's a lot of design space still out there to explore.
The top-down centralization of mid-20th century technologies wasn't caused by "money," it was caused by their technological nature. They inherently relied on limited radio spectrum that was doled out to specific companies by the government.
Yes. I thought maybe Blockchain might be able to provide some of that robustness for the public web, but it's all about speculative investment now.
A fully traceable public ledger is kind of a wet dream for an intelligence agency/evil corp. Funny how much faith we put into that, and look at it now.
What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
And even when TV was limited to a small few national OTA (over the air) channels, different people would tune in at different times to watch their different preference of shows. So just because content was fewer, it didn’t necessarily equate to different people consuming the same content.
The real crux of the modern problem isn’t preference drive choices. It’s independent publishers being driven out by corporate greed. Eg AI traffic making it unviable for independent blogs. But then one could argue that the current ecosystem, where the barrier for publishing being so low, is an anomaly because historically that was always prohibitively expensive. To go back to the TV example: you couldn’t commission a show without deep pockets and a lot of TV exec contacts.
I’d love our golden age of information to be persistent. But, and as yourself and others have alluded to, that’s not guaranteed unless we fight to keep it that way.
>What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
Exactly. The shared reality is the experience. You both were given the same opportunity to take in the same information. The different opinion is what comes out of the other side of the meat suit processing the experience.
Algo driven results changes the experience by changing what the user perceives. So there might never be a shared experience.
“Never” is a little dramatic. Sticking a TV show on from a streaming platform and watching it as family is still less effort than going to the cinema.
> For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
yeah because there is a shared reality in which you can both see the same movie. enjoy it while it lasts.
I think of it more as a basis for shared culture - I could chat about Seinfeld / Friends, and odds were high that the other person would know what I was talking about.
I imagine in prior ages this was perhaps more about folks reading a core collection of books. I expect there was less shared culture than in ^
Now it seems like we're in an era with (potentially) less of a shared cultural bases than either of those times. I've certainly met people where there's really no overlap - their core beliefs, values, etc - fundamentally different foundations. If that becomes more widespread, it's hard to know what that will do to our societies.
The thing is, people I’m friends and thus likeminded, we watch the same shows and still have those kinds of conversations. And people I wasn’t likeminded with never used to watch the same shows, even pre-streaming era, so I’d never had those conversations.
Take your examples, Friends and Frasier were aired at roughly the same time in the UK (obviously on competing stations). Most people watched Friends, but that show wasn’t really my thing. So I watched Frasier. As did my closest friends. And we’d chat about that. But I couldn’t join in with work colleagues with Friends chats.
It’s a bit like sports matches. People at work will talk about football/soccer but I’d be more interested in Snooker. So I’d chat snooker with friends and duck out of the football work chat.
Independent blogs are perfectly viable. It is basically free to publish. What's not viable and never was is to be a professional blogger. At best, some people could try to be professional shills with blogs as bait (it is good if these go away; they were noise), and an extremely tiny minority might actually make a living with people paying to support their blog directly.
I think being a professional blogger is on the rise, the driver for the viability of being a pro blogger has risen. Substack/medium etc. streamline the getting paid part. The fracturing social narrative, (rise in censorship, cancel culture, tribalism) actually drive people to pay to gain access to periodtical writing from specific people seen as authorative on one or more subjects. These subjects that mainstream media are too shy to surface, too dry to feature, or have a tendency to twist to a narrative that supports the sponsors etc.
> What even is a “shared reality”? For example my mum and I could watch the exact same movie and showing at the cinema and still come away with a completely different opinion of it.
A: It excludes opinion.
Everything is opinion. A shared reality is just where opinions coincide.
The shared reality in that example includes only facts: OP and mum went to a movie. The movie was called X. The actors were Y and Z.
The original point of this subthread was that we can't even agree on a shared set of facts anymore! The political Right has an entirely different set of facts than the political Left, and when our reality is derived from those facts, we cannot share the same reality anymore.
If the facts can differ then they’re not facts. What you’re then dealing with is propaganda.
Then, we're just back to OP's conundrum: "Everything is opinion." I'll say grass is green and you'll say grass is red, and there can be no fact because they differ.
Not really no. Some things can be proven while some cannot.
The colour of grass is something that’s scientifically measurable.
My point is that some broadcasters are very strict about sticking to the facts. Some will misrepresent those facts but not technically lie (eg discuss statistics that favour their partisan view but don’t share the full context behind those stats).
And then there’s broadcasters like Fox News that promote flat out lies. It’s often not even an interpretation of the truth. It’s shared as an opinion but the evidence clearly proves it’s just lies presented as facts.
And I really wish there were criminal charges for broadcasting blatant lies. Not just in America, but most of the developed world where populist-propaganda is used to brainwash gullible voters.
Communication cannot be scientifically measured. Someone's expression being a "blatant lie" can only ever be an opinion, fundamentally.
Broadcasters absolutely have a responsibility to ensure they communicate truthful statements in exactly the same way that any other profession has a responsibility to ensure their output is correct.
Everything is perceived, and our perception gets filtered based on our prior experience. Just because we're not directly perceiving objective reality, that does not mean _everything_ is opinion.
JFK was shot, that's not just my opinion.
Unless you choose to believe some kind of post-modern crazyness, in which case - have at it.
You don’t actually know JFK was shot from your own personal experience. You just choose to believe what credible source claim. And even if you were there, who’s to say that your fallible memory is incorrect?
To be clear, I don’t think this line of thinking helps much. But it’s an interesting philosophical dilemma.
"Who’s to say that your fallible memory is incorrect" - you can very easily test if someone has a bullet through their head. Well, assuming someone didn't go and steal their brain.
I'm saying that not everything is opinion, that there does exist a category called 'fact'. And you're talking about something else - a chain of trust with respect to those facts.
This is starting to go down the crazyness line I had mentioned. I get it's fun to philosophize about such things, but it's not how we live.
> you can very easily test if someone has a bullet through their head. Well, assuming someone didn't go and steal their brain.
Do you have access to their head? How do you know that someone didn’t swap that head for someone else who had been shot?
Like I said, this is purely a philosophic argument. I’m not actually trying to suggest that there isn’t such thing as “facts”.
> And you're talking about something else - a chain of trust with respect to those facts.
No, I’m making a psychological point that our perceptions are what we use to construct our reality, and that our perceptions are malleable. If they weren’t, then drugs would have no effect, psychosis wouldn’t be a thing, and people’s opinions wouldn’t differ about even the most straightforward things like god, globe vs flat earth, and so on and so forth.
> This is starting to go down the crazyness line I had mentioned. I get it's fun to philosophize about such things, but it's not how we live.
I agree. As I said in the previous comment:
“To be clear, I don’t think this line of thinking helps much. But it’s an interesting philosophical dilemma.”
There is an pink elephant behind you. In practice, I believe that at least after the dust have settled, events of the type of "Count Ferdinand was shoot dead infront of 100 witnesses", like maybe 1/50 are fake? Dunno what share.
The faking part would be in who is blamed and what country to bomb etc.
Not necessarily. A shared reality is one where people are discussing the same thing (within a reasonable margin) even if filtered by their individual perceptions, and drawing different conclusions.
I don't know the word for the opposite of shared reality (maybe private individualized realities), but it results in a scenario where you have nothing in common with other people other than "I saw some thing" and therefore it becomes very difficult to discuss anything with them.
"I don't know the word for the opposite of shared reality"
Different cultures?
Maybe, but not necessarily.
I think different cultures can share the same reality, though filtered through different lenses.
If two people from different cultures watch the same movie, one finds it ok and the other offensive and disrespectful, that's a shared reality filtered through different cultural backgrounds. Even if either version of the movie is dubbed, censored, or altered in any way, there's still enough common or shared material that a discussion can be held.
If two people had completely different experiences and watched completely different, tailor-made movies, then there's very little for them to discuss or even disagree with. "I haven't watched your movie", "I haven't watched yours either", "Mine was better", "I guess... I wouldn't know, I thought mine was cool", etc. No shared reality.
If everything is opinion then there can be no shared reality. The only "reality" that exists is your opinion about there being one.
Like whether you teo were even in the same room watching that movie.
Opinion and preference are the same thing. And thus personalised search results, choice of shows in the streaming era, and so on, are just a result of that personal bias.
Well, loss of (control over) shared narrative was one of complains of the Church against the printing press ...
And institutional control of information didnt really start again until the Statute of Anne struck up a deal between printers and the state giving printers copyright over written works and the state the right to cencor and ban works the disapproved of nearly 250 years latter.
Radio gave us Demagogues like Father Coughlin in the US and Hitler in Germany, and we still havent dealt with the problems radio introduced when television came along and now we have the internet with social media, podcast, and algorithmic echochambers, and algorithm induced radicalization. Radio television internet any one is just as much a shock to society as print was and we still havent figured out how to adapt yet to any fully.
The VCR became widespread during the 1980s
> The industry boomed in the 1980s as more and more customers bought VCRs. By 1982, 10% of households in the United Kingdom owned a VCR. The figure reached 30% in 1985 and by the end of the decade well over half of British homes owned a VCR.
Another good one.
My impression of the VCR is not too many people managed to use it for DVR-style rescheduling live TV. Although it did have that feature. We mainly used it to watch movies that we had already been in the theater.
But, for sure, another fracturing. We weren't all talking about the nine movies in the theaters, but whatever we had rented and watched on tape instead.
My recollection was that the VCR's main selling point was the ability to record TV and play it back later, and I can say anecdotally, that that was the primary way my family used ours in the early 80s. Only later did the idea of Blockbuster and "using VHS merely as content distribution" take hold.
My family used it to capture programming we couldn't normally be around for, especially late night BBC content on PBS. Pretty much all we used it for was DVR style recording.
My neighbor's house was awash in PBS recordings! https://www.pbs.org/wgbh/masterpiece/ https://www.pbs.org/wgbh/masterpiece/ https://www.pbs.org/wgbh/masterpiece/ https://www.pbs.org/wgbh/masterpiece/ https://www.pbs.org/wgbh/masterpiece/ https://www.pbs.org/wgbh/masterpiece/
> But as with any "beginning" of a long, diffuse process, people might quibble about the specifics.
Yeah, absolutely. But for the web, at least for my daily use of it, this was the first big "shift". Learning how something was spelled by someone googling it and quoting the number of results to the others.. good times. And you can no longer find out what the page with highest page rank says about "pizza" regardless of where YOU happen to be.
Yes, there's certainly less of a monoculture now that people have greater choice in what media they consume. Personally I have trouble seeing that as a bad thing. Greater diversity of taste, culture, and opinion makes for a richer, more interesting society in my opinion. Though I recognize it does have downsides.
I wonder if there's a way to keep the positive effects of such diversity but mitigate the drawbacks. Certainly echo chambers are undesirable.
Culture is shared experiences. There is no culture if everyone lives in their own little bubble of reality.
There are other people in those bubbles, it's not an isolated void. HN is one such bubble. Yes it's a smaller bubble, there's not nearly as many people here as there were watching broadcast television back in the 1980s, but it would be wrong to say there is no culture here. Same goes for the millions of other small communities out on the web and social media. But as I said, it's certainly true there are downsides associated with that diversity.
> I wonder if there's a way to keep the positive effects of such diversity but mitigate the drawbacks
The answer society mostly settled on is representative democracy.
You appoint qualified candidates via consensus, and then give them the microphone and some measure of control as long as the consensus stays strong. This is sort of what we have now with influencer culture.
The flaw has always been human nature. It turns out that half of all people are of below average intelligence and they consequently make bad decisions. Especially if the smarter half use them as pawns to assume control, which is what US politics has become.
Once the Morlocks and the Eloi have settled into their roles the entire thing falls apart and democracy dies screaming.
> It turns out that half of all people are of below average intelligence and they consequently make bad decisions.
Worse than that; it turns out that even people of above average intelligence consistently make bad decisions! AND they still attempt to control others!
Good point. Humans in general are the problem... ChatGPT for president!
All of those are pretty specific to US national culture whereas GP is talking about something more global.
Good point! Interested to hear about the RoW. Where are you from? Was there pre-custom search results fracturing or did you have a (relatively) uniform national media before that?
I never thought of the DVR as a particularly significant development. I thought of it as just a replacement for the VCR. That's what it was for me. But perhaps in some areas DVRs were used to a greater extent than VCRs had been?
Tivo made it very easy and drove the DVR thing IMO. The whole "programming a vcr" was a meme back in the day for something overly confusing.
Better quality recording than tapes, and easier (for many people) to program.
Yeah, I readily concede it's lesser. For me it was the moment that we stopped watching a particular thing and discussing it the next day. But maybe that's more personal biography than sociology.
I have a dream, to ban targeted advertising globally. It would solve many problems and possibly, cherry on top, kill some corporations
I've always thought that if online advertising as a business model had been entirely banned from the internet in the 90s, the internet would've looked like a much better place today.
And micropayments would've taken off in a big way.
The only reason why micropayments doesn't really exist today isn't due to technical reasons, it's because advertising is too dominant, and because consumers got trained on expecting everything on the internet to be "free", not realizing they're the product.
Stop the purchase and sale of user data, period. Kill the weed at the root.
There are many indirect ways to get information though. Better would be to ban "discrimination" of customers/humans (no loyalty programs for long term customers/whales or people who know people), and ban of usage of all personal information for decision making except short pre approved lists. Most industries would do fine with "payment received", some maybe could make legitimate cases for also "delivery address" or perhaps "age". Would probably also be good for efficient competition on merit, which should appeal to the groups with enough wealth to use it to accumulate wealth.
Online advertising is what drives the desire for user data. Ban online advertising and the demand for user data withers.
A person can choose to ignore or avoid ads, but having their data bought and sold is solidly out of that person's control.
Who's selling user data? Do you think Facebook is out there selling it? How do I buy it if so?
not directly, but https://brightdata.com/products/datasets/facebook
Or, just ban advertising.
Humans are only now figuring out pointy shoes are bad for people's feet. This species is wrong about almost everything when it comes to its own well being and like pointy shoes - it's a completely unnecessary self-own that barely benefitted anyone.
> ban advertising
I think advertising runs a little deeper into our identity and history than you are thinking. If you have a sign in your window describing your wares, that's advertising. If you're standing on a street corner inviting customers to enter, that's advertising. If you're doing nearly anything to make the world know that your business exists, you're advertising. I'd rather not live in a world where "ad cops" see fit to intrude and opine on such activities. Attacking PII brokers, on the other hand, would do far less collateral damage.
I'm with you on the shoes though.
When I express a slogan, I am fully aware that it lacks nuance and is not to be taken literally.
Pedantry is not a counter argument, it's a deflection. We are not in court.
Personalization came along in MMOs around the same time, and it was hard to communicate the difference to people between a persistent MMO zone as a "place" where people and things exist, even if those things are inappropriate for your character or hostile to the point that you can't survive there or too crowded for the intended flow... versus an instanced MMO zone, a little pocket universe for you specifically that has no strangers in it, which is more of an "experience" or "cut-scene" for you to imbibe prepared content.
A heavily instanced MMO may as well be a single-player or co-op game with a shared chatroom as a lobby, there is absolutely none of the emergent play, none of the characterization of players & groups of players, nothing original. Instead, give me the same ground to walk on as everybody else, even if our tread wears all the grass off of it; Half of my life in these games was getting bored waiting for a mob to spawn and fucking around making friends waiting in line, or ganging up on the guy who wanted to cut in line, or waging war on the clan that currently controls the needed territory, or grouping together to fight mobs higher than I should be able to hit (which the instance will just Adjust Downwards). Instancing and other auto-personalization algorithms cut the annoyances that are central to human socialization, achievement loops, and wayfinding.
Similar impact to non-massive multiplayer online games. Skill based matchmaking destroyed the third place that once existed on public game servers.
I still don't understand the appeal of an online game where you're algorithmically guaranteed a ~50% win rate after a calibration period. Seems to be selling well enough though.
I think the lure with matchmaking was that when it was introduced, it made you really good compared to before by easier access to better practice.
In Data 2 the skill level shifted upwards enourmously for pro players and I feel also for normal players. But the scene also got kinda lame and boring with too optimized strategies.
And in the process I think it alienated all the players for whom getting really good wasn't the goal.
I never wanted to be the best CS:S player, I wanted to hang out with the regulars on the server I was a regular at. Like a bowling league or a group bike ride.
> I tried to explain why personalized search results are the beginning of the end of a shared reality
It ignores the fact that there was no such shared reality. Just go and look for the time when everyone agreed if racism existed.
Funny how you go on about the world ending while merely repeating it:
https://en.wikipedia.org/wiki/Seduction_of_the_Innocent
https://en.wikipedia.org/wiki/Parental_Advisory
https://www.bbc.com/news/magazine-26328105
What's more likely? Reality is collapsing or you brainwave synced with elders riddled by post war and cold war PTSD and who huffed lead gas smog: https://www.scientificamerican.com/article/brain-waves-synch...
Have watched every generation act like this since getting online in the late 1980s; act like the end of their social truisms are akin to a black hole forming in the middle of the sun.
No; it's just them realizing they have fewer days ahead than behind.
Before search engines, we had personalized search results in the form of yellow pages. The yellow pages in Chattanooga are different than the Wichita.
There was no global yellow pages where plumbers from Odessa were listed beside plumbers from Tokyo, was there?
If you were looking for something more esoteric, you were going to find different results in the World Christian Encyclopedia than you would in the Encyclopedia Britannica
Wait, so anything that's not sold everywhere is personalized? Don't think so.
(At the playground)
"...but how do I know that what I perceive as a certain meme, you also perceive as that same meme??"
I think a lot of this is "accelerating trends."
That is, the www was "being killed" already in a lot of ways. "Democratization" and suchlike open ideals had already been receding for a long time.
In part, this is because of "democratization" of access. The old web was a self selected, subpopulation. After smartphones, it is the whole population.
Facebook's walled garden and other apps using the web as "merely infrastructure." Algorithic recommendations replace hyperlinks, until it no longer "a web."
Google, imo, represents a sort of intermediate stage. "Pagerank" the original tech behind their search relied on the hyperlink web structure... and also degraded it.
That was SEO paradigm. Now there is the LLM equivalent of SEO.
So... LLMs eat the web, while also making it redundant. The previous paradigm and all it's contents whill exist as a ghost within the new one.
I always hated the term "democratization" because last I checked there was never any consensus on what should be done, always MBA pricks shoving whatever they wanted down our throats.
If the timeline of things as you present them is taken as accurate, then I would say the www being killed was good, with FB opening up to everyone being the inflection point toward bad.
Webrings are nostalgic, but the old web sucked, and there was a dearth of information and fun to be had. I remember getting online and being able to confirm "nope, no new anime content worth looking at today" by quickly checking Usenet and Yahoo's index (which was updated manually) and a few webrings.
now, I'm able to watch a show that aired in Brazil the day after, with English subtitles.
People who long for the web of the 90s, I always feel like they must not be very interesting. Imagine complaining that we can't go back to libraries of the 1970s when the ones today have 3D printers, DVDs to rent, etc. because "there are too many teenagers"
That's confusing content access with monopolised algorithmic promotion/curation of belief systems, using selected news and opinion pieces to modify beliefs and behaviour.
It's nice you can view (pirate) whatever content you want, but not so nice that algorithmic platforms pretend to be neutral social spaces when in fact they're being used to promote certain political and cultural beliefs while suppressing others.
I don't know about anyone else, but anybody who goes into a "1970s library" (I don't know what that is, but I guess it means a good old-fashioned library with like, books) and complains about the lack of 3D printers or that their obscure animated series catalog isn't wide enough would embody the definition of an uninteresting person to me.
I am one of the people that long for the old internet. I don’t long for it because there are too many normal people on it now.
The old internet was interesting, open, simple, un-gated, and relatively egalitarian with few corporate interests in commanding positions. The new internet is basically a predatory environment disguised as a library.
I think when people say they miss the web of the 90s what they're really missing from the 90s is themselves.
If you judge the internet and the world by "no new anime content today" things have actually improved. But some would think that's not a healthy way to see things.
> People who long for the web of the 90s, I always feel like they must not be very interesting.
Judged on an anime freshness scale, probably not.
> People who long for the web of the 90s, I always feel like they must not be very interesting. Imagine complaining that we can't go back to libraries of the 1970s when the ones today have 3D printers, DVDs to rent, etc. because "there are too many teenagers"
The Internet used to be a space for high-IQ people, and now it's filled with low attention span, low impulse control people. Tik Tok and Instagram are the worst offenders in this regard, because it enabled borderline illiterate people to communicate on the Internet.
I think we still have many high-IQ people on the web. Its just they are not making tiktoks. However, they are harder to find since the pool got bigger.
The ease of use introduced the average individuals to the web and online gaming. When I was younger and online gaming was gaining traction my "clan" hosted servers. This is how I started dipping my toes into programming and server configuration. Literally, everyone that was in the clan was quite intelligent. We had debates and discussions on IRC, forums and teamspeak.
Now, gaming and the web is just trash full of brain rotted individuals. Communities have been destroyed by lack of self hosted or self owned servers. Forums have been destroyed by discord. Games are being destroyed by microtransactions.
A quick read of any old usenet arguments should instantly disabuse you of this absurd notion. Or the old archives of stupid IRC conversations.
The old internet was a space for people who were incentivized to spend significant effort dealing with technical things to connect to strangers, which has nothing to do with IQ or "intelligence" or anything like that.
There has always been plenty of people who are extremely stupid, but have no problem following technical minutia as required for things like the early internet.
Remember that early access was extremely expensive. Usually it was billed hourly, at significant rates. This means that the primary filter for internet users back then was connection to resources, through an academic connection, a corporate connection, a government connection, or a lot of personal wealth.
The quality of internet users early on, if it even existed (which I dispute), was more driven by organizations like the government, academia, and defense contractors.
And as much as people cried about "Eternal September", the internet mostly survived as a distributed place with real communities until the late 2000s, where it started to be really profitable to build walled gardens for Advertising economies.
The walls didn't really come up until the 2010s: Facebook being a dominant platform, Google owning Youtube and leveraging that to try and force their own social network, Apple controlling most phones, old forums dying, and the rise of AWS to centralize internet software design.
Being able to deal with “technical minutiae” is highly linked to intelligence. What you’re referring to is lacking social graces, which is different than lacking intelligence. Having a hobby that involved reasoning about abstract things like IRQ assignments is a pretty good proxy for IQ.
Similarly, back when computers were expensive, people on the internet generally were higher income, which is also quite significantly correlated with intelligence.
We’re burning the new library of Alexandria on the altar to a new god.
Worse, we're replacing the contents of the library with shit and leaving it standing. Burning it down might be the correct response.
It was already on fire, AI just pointed a leafblower at the coals.
Before slop farms started using ChatGPT, they were using outsourced writers who had a tenuous (at best) grasp on the language and subject. If you searched anything on the web, your first page of search results was generally incredibly fluffy articles.
For example, you would search for "How to use python async" or whatever, and every non-stackoverflow result would go something like "Python is a useful language for async library usage. Python is a language created in... You can install Python by... Synchronous programming is... Asynchronous programming is... Python synchronous code looks like... Python async module code looks like... Popular libraries for async are... Other languages have async such as... To import async module you can..." with every paragraph interspersed with an ad. The content was surface level, relying upon blatant plagiarism of better-written articles and forum posts without attribution. It was an utterly miserable experience if you placed any value on your time.
The only thing that's changed is the quantity of slop, and the karmic irony that the content farmers who put the writing profession into a race-to-the-bottom were themselves discarded, in much the same way as the scabs who replaced striking workers were frequently discarded in the industrial era.
And that's without even getting into the subject of the people who made money as Internet point farmers of Reddit or Twitter, who would repost generic content then sell their accounts to spammers.
The Internet was already going in this direction; the only difference is the rate at which it enshittified. I agree that it's in a worse state than it was 10 years ago, but let's not look at the past through rose-tinted glasses.
> The only thing that's changed is the quantity of slop
The most important thing that's changed is the proportion of slop.
When it was humans originators v human sloppers, we stood a chance.
Now it is going fewer human originators v. vastly more machine sloppers. No chance.
True, LinkedIn posts were AI slop before AI was AI slop
Worse, we gave the matches to the new god then stood back and watched.
Well, I called it in 2015 because they were already adding features to sites to circumvent archiving and other things. Forget AI. That has nothing to do with why the collective memory is being lost.
I find it so interesting how many comments about AI start with a disclaimer that of course it's very useful, before pointing out a negative of any sort.
I do it too so it's not a criticism, but I find it kinda interesting culturally. It's a reflection I think of the environment that we've created for ourselves, or has been created for us (or some combo) where somehow you can't point out an unqualified negative of this tech.
This to me is actually one of the clearest bubble signals, just because if it really were as good as all that, it'd be completely redundant to remind everyone of that when criticising it.
I think it might be a sort of weird collective psychological thing where we must not admit what seems pretty clear, just for fear of breaking from norms.
But, of course, yeah it's very useful for some stuff and I use it all the time :-)
Part of the problem is that there is so much overlap between GenAI's biggest boosters and people who reflexively welcome on any tech change as an unalloyed good. People can be really fanatical about this stuff. Any nuanced discussion of tradeoffs or negative externalities draws insults of being a Luddite.
My own general thoughts are that AI is extremely useful in discrete cases, will cause an immense amount of cultural and societal harm overall, and there's no clear political path to a healthier development trajectory. I don't think that's an uncommon stance, but people espousing something like that get a lot of hate in a lot of discussions.
I can easily imagine an arms control style series of agreements between the handful of leading AI companies and countries which are imperfect but still quite effective at mitigating the harms without stifling the upsides, but it seems clear that the current leading actors in this space have zero interest in that.
It's a very wide technology, there's going to be use cases where it works well. I think without the "of course it's useful" then you'd just be flooded with examples of where it works well as counterpoints whether they're relevant or not. Critics have to set the stage that they're criticizing a particular use case or aspect of the technology and not the whole thing all at once.
i think its cognitive dissonance. everyone is using it pretty frequently and so to criticize it creates this tension like "if its not actually worthwhile then why am i using it so much?" and the most prolific response to that dilemma is to write it off as saying "well of course its useful in many ways" ... with the unspoken second half of that sentence being "it's useful in the ways i use it"
It happens everywhere with every topic on the internet. You need to let your fellow tribemates know you're in the same tribe before you say anything negative, otherwise they will discard your argument and flame you to death.
Go into any subreddit and you see the exact same pattern.
More than a bubble signal, the “but it is so useful” disclaimer is what makes it a tragedy of the commons.
“I know it is ruining all that I hold dear, but what else can I do but profit from it like everyone else?”
To be fair that is also true to the Internet itself, it's very useful but has negative side-effects.
but when you criticise it you don't feel obliged to point out it's of course useful. This is just obvious unstated.
I think you're describing political commentary habits becoming general writing habits.
It works with just about anyone but try and say something positive about Trump and, if you don't add a qualifier like that, the first reply is going to be something like, "I guess you like ICE murdering babies." And then the conversation you wanted to have gets completely derailed.
The majority of people's opinions on the subtopics of complex or controversial topics have to align with their opinion on the overall topic.
AI is bad so it must both be useless and be using up all the water.
Well I think theres a lot of social pressure to disclaim "yes of course its useful" because if all you do is say "hey AI has this problem," you'll have AI bros replying to you telling you about all the good and saying you're misrepresenting AI.
I agree and I think there is also a more insidious meta aspect: for every comment like this there is the worry that some X multiple more people hold that view but won't say it.
Everyone else also thinks like that and so isn't normally unambiguously critical, so you end up not knowing who really thinks that way and who's just doing it because of the game theory.
It could even be practically nobody. There's no way to tell until the deadlock breaks.
> I find it so interesting how many comments about AI start with a disclaimer that of course it's very useful, before pointing out a negative of any sort
People have to post very defensively online, otherwise some "um ackshully" guy will be immediately jumping down their throat
If you write anything negative about AI without stroking the ego of the AI bros first by praising their machine best friends they will be all over you calling you a luddite
Yes, this is really stereotypical HN: You can't really post an unqualified generalization here, because someone will inevitably jump out of the woodwork with a canned "You said X is generally Y, but here is counterexample X0 which is not Y! Your post is incorrect!"
Yeah. It's all over on HN for sure. I associate it a lot with Reddit too
Upvote-based systems and their consequences have been a disaster for online discourse
I called this in 1994 when that Lee dude opened the floodgates to everyone and everyone brought their problems to the web.
how I handle this - if I find an interesting web page - I save to pdf & save to my external sdd. ideally I should get one of the 8tb HDD to save material I see on the web.
coz yeah good quality material is disappearing on the web fast. the only thing remaining is people building their own private search indexes.
Not that this does not work, but there are services for that you could run. If you got (or plan to) a homelab, that is a nice add
I've been using Zotero for that, although it gets slow as hell eventually because it is a whole-ass firefox instance. They're real references so I might as well treat them like real references.
I've been working on a project that basically is a superset of Zotero, just efficient. As if you expected to have more than a few thousand references. I'm sure I'll get it done after the web is long gone.
The writing has been on the wall for quite a while, thanks to Google itself and SEO, plus Google advertising channels.
Some websites already had minimum “for humans” content for years, as it was all used to appease the algorithm.
And Social Media was the cherry on top, with information overloading, walled gardens, personalization and companies using it to pull a “Hey Fellow Kids” when using it for marketing.
Very soon, everything easily accessible on the Internet will be a never-ending loop of AI slop. Will this push us back to analogue things? Not in an apocalyptic scenario, but in general, and by analogue, I don't mean no use of computers but use of computing with a firm and predictable human touch and control.
Will there be venues left (largely) unencumbered by AI to even turn analogue? The top-down control of our industries, financial systems, education, etc by a few mega corporations and fewer mega rich people, who also have vested interests in the advent of AI as AI corps are trying to be, means there's nothing left where AI is not inserted in every vein and nerve and nerve centre.
It is as if a few people in the world are trying to turn this world into something Frankensteinian because they think that then they will get the chance to be the only ones to control this Frankensteinian. They might as well succeed. To what end? I do not believe even they know that. Blindness of greed.
what value do we get from AI?
Last month I rebuilt my pool's plumbing system with the help of YouTube and GPT. I was quoted $6k and it ended up costing me about $1k of materials and a few days of work. This was my first exposure to any kind of outdoor plumbing, I'd never touched PVC before.
It's unlikely I would have had the confidence to do it from YouTube alone, specific diagnostic help and a full diagram to work off of was extremely helpful. I had it prepare an SVG of the whole assembly with all the measurements and parts labeled.
I'm glad it went well for you. But I run a not-for-profit pinball museum, and I have specifically had to ban volunteer repairpeople from using the chatbots because it will confidently propose idiotically wrong solutions. Which will then get proposed to me, or just directly applied, with equal confidence. At this point I'd rather hear, "My horoscope said..." than "ChatGPT said..."
I expect that it's good for common use cases that's well documented. But then, so are a lot of other approaches.
> I run a not-for-profit pinball museum
Yeah, I would imagine pinball machines are at least an order of magnitude more complicated to maintain than above-ground pools are. That situation sounds really annoying, I feel for you.
For sure, but I think the deeper problems are how much the industry changed over the decades and the extent to which pinball repair content comes from amateurs opining on forums.
Pool technology has been more stable, and there are more people out there writing well-informed content for LLMs to extract and present as their own.
Two years ago it would have been insane to say that you got help from ChatGPT to fix your outdoor plumbing, people here would have been frothing at the mouth for merely suggesting it. Four years ago it wasn't even on the radar of future possibilities.
The quality it is today, is the worst it's ever going to be.
Maybe? It's a plausible theory. But these operations are all wildly unsustainable financially at the moment, and it's not clear where they'll get future data from, having destroyed a lot of the incentives that generated their current source content.
Bubbles are not a great time to form intuitions. WebVan [1] and Kozmo [2] also seemed to herald a new age. Decades later, brick-and-mortar grocery stores and convenience stores are still doing fine.
[1] https://en.wikipedia.org/wiki/Webvan
[2] https://en.wikipedia.org/wiki/Kozmo.com
We are in the golden age of AI. Things are going to get worse at some point. It happened to google, it will happen to AI.
They already have all the data, all the money and all the chips. It’s actually the best it’ll ever be, as newer models will have to start paying off all that capex, newer models will be trained mostly on slop, and the SEO and influence op leeches will have begun their arms race to insert their products and values into the training data. We’ve seen this pattern before. AirBnB, Uber, social media, streaming video, et cetera didn’t get better once the VC money ran out and they needed to start turning a profit, they got much much worse.
By chance is it the one in Asheville, NC?
That's a good one! But no, we're in Chicago: https://www.theflip.museum/
The horoscope reference is resonating for me.
I know very little about pinball machines, out of curiosity, do some come with any amount of schematics or diagrams?
I have schematics as early as 1937, but the odds that the manuals stayed with the machine are pretty low. (Usually these machines were owned by somebody who had a bunch, and I suspect the manuals tended to get collected centrally.) You can generally buy manuals on eBay, and some are being reprinted by people who own the original IP.
The internet has a lot of scans of wildly varying quality. The AI industry's hunger for data means there has been a lot of progress in OCR and data extraction, so I have notions of taking something like PaddleOCR and trying to turn the scans into a cohesive reference site that will work well on phones, etc. But I'm not sure when I will get to it.
If people out there are interested in working on a project like that, let me know. Email me at william@theflip.museum.
Not OP, but this channel does a lot of walkthroughs of popular machines and explains how those systems work in-game:
https://youtube.com/@papapinball
I also strongly recommend the Technology Connections series on how electromechanical pinball works:
https://www.youtube.com/playlist?list=PLv0jwu7G_DFVAUoqtVxFV...
I would love to put them on a kiosk in the museum.
> This was my first exposure to any kind of outdoor plumbing, I'd never touched PVC before.
I find it telling that the highest praise for LLMs comes from people using it for something where they admittedly have very little domain knowledge. Domain experts usually mention major caveats. I've been testing them on subjects where I already understand the problem well, and I've yet to see any outputs that would make me trust them on things I don't already know.
https://en.wikipedia.org/wiki/Gell-Mann_Amnesia
https://en.wiktionary.org/wiki/confidence_trick
When I apply LLMs to domains where I'm already an expert, I find it particularly lacking when I want it to do deep, difficult, novel work that requires precision. On the surface, the output looks pretty amazing at first. But when I turn a critical eye to every detail, I end up finding a lot of flawed "thinking", and the lengthy process of fully understanding what it generated and cleaning it up to my standards makes me question the entire value proposition. However, I find it does a great job being a low-level automaton sort of assistant.
For instance, in the domain of software engineering: I would not trust it to implement a major architectural change, or a groundbreaking, complex new feature. I would trust it more (but not completely) on something like a refactoring that may touch thousands of lines in a fairly mechanistic way, but that was a little too-complicated for simpler tools likes regexes. While that's kind of a nifty use, I think it's fair to say that non-LLM software purpose-built for such tasks can probably do the same thing more effectively for less real cost (meaning the currently-subsidized real cost of all the training and inference power burn, etc)
The greatest living mathematician uses it for Math
https://siliconreckoner.substack.com/p/terence-tao-on-machin...
To me it seems like people who are confident in their expertise generally find it useful, even if it's imperfect
That's a good point - I've heard people make the same one whenever LLMs are brought up. I hope someday you're able to get more value from them.
Anyway, my pool's looking great and I gained some new skills. I probably could have gotten there with books and YouTube alone, but having another tool at my disposal made me a bit more confident.
I remember this was brought up about a month or so on HN, in the context of describing how one person's experience using LLMs can be so vastly different than another person's:
"LLMs seem good at things you are not good at."
So, if you've never touched PVC before, LLM sounds plausibly competent--it may actually be or it may not be, but you'll walk away from it thinking you learned something. If you are a professional plumber and ask an LLM the same thing, the output will more look flawed and possibly dangerous.
Same for software writing: If you're not a good software developer, you probably think an LLM is great and writes much better code faster than a human developer can. But if you are a good software developer, LLM output is slop and requires huge rework to be passable.
> So, if you've never touched PVC before, LLM sounds plausibly competent
well to be fair, it's sounds about as competent as your average homedepot employee. He's wasn't doing something super complicated, cutting and gluing PVC for an above ground pool is very common and doesn't require a plumber. I used some youtube videos to fix my dishwasher, i didn't need a professional service agent from the manufacturer I just needed some pointers.
as for software writing, for standard everyday enterprise app work which is typically just CRUD and moving data around it works fine. That kind of software does not need to be a highly tuned work of art to meet the requirements.
LLMs seem good at things you are not good at... primarily because you lack the skill to actually judge the goodness of their work, and they present their work confidently with an air of authority.
This sounds accurate.
I am not a mechanic by trade, but I know nearly everything there is to know about working on an ICE car. I did paint cars professionally for a bit.
LLMs are absolutely full of garbage advice, when I try to use it for troubleshooting. However, people who don't know anything about cars are telling me it helped them fix issues. I am assuming their issues were maybe surface level and something I would just know without even looking at any manuals, because when I use it for complex problems it just doesn't work for me.
if it helped them fix issues then it helped them fix issues. Good for them. Maybe their problems didn't rise to your bar but at least they were able to get it solved. That's very useful and empowering to people without the direct knowledge and experience themselves.
That's the entire problem. The main thing LLMs gave you was confidence, and the thing to understand is that the confidence an LLM gives you is often utterly false and baseless.
Why were you not confident with literal how to videos and documentation, but became confident when a chatbot generated probable text?
LLMs are infecting us all with Nobelitis.
Well, I guess another plausible explanation is that usually "domain experts" are people that get paid for their expertise, and have a vested interest in saying that LLMs cannot replicate what they want people to pay them for.
It’s likely it missed some crucial subtlety that will come to bite you in the ass down the line.
That’s the thing with AI - its responses sound plausible enough to non-experts but time and time again I see experts in any given field being able to identify AI content by pinpointing subtle but crucial errors. That’s one of its dangers - it gives you enough confidence to shoot yourself in the foot.
Human workers also wreck these jobs horribly that come to bite your ass in the end. I think an intelligent person equipped with AI and common sense and real stake in the thing being well done (because it's their own, so they care) is better than whatever is possible to pay for or book in a realistic timeline from another human.
Also, ask people in the trades to review each other's jobs. They will harshly criticize each other too for missing basic things and then go on to vehemently disagree. As an outsider it doesn't mean much that an expert found some fault. They always find something to nitpick.
> Human workers also wreck these jobs horribly
Yes, but with a human worker there is a chain of accountability. With AI, there is none.
Have you had an argument or dispute with a contractor before? It's better to just do it yourself when you can, and fix it when it breaks
> It’s likely it missed some crucial subtlety that will come to bite you in the ass down the line.
Maybe, but that's part of the experience of learning. I plan to maintain pools for the rest of my life, if I made an oversight which costs me down the line then the lesson will be that much more memorable.
This is an above-ground pool with a pump and a filter, the stakes are relatively low. In the absolute worst case I could rip it all out and pay a pro to do it for the price I was quoted.
If that's a pool for human use, the worst case is a bacterial/fungal infection due to incorrect filtering.
A pool filter is for removing physical debris, it's not really for preventing bacterial or fungal infections.
I guess what you're talking about is sanitizer, in my case we use chlorine. I test it every time we swim, but I wouldn't have needed an LLM for that. It's very straight forward to maintain pool chlorine, my Dad taught me that when I was 13.
> It's unlikely I would have had the confidence to do it from YouTube alone, specific diagnostic help and a full diagram to work off of was extremely helpful.
People have been DIYing swimming pools for decades with the help of… books:
* https://www.amazon.com/COMPLETE-GUIDE-SWIMMING-CONSTRUCTION-...
Where do you think ChatGPT got its information from?
> It's unlikely I would have had the confidence
Perhaps with good reason?
In the 90s and prior you could have done the same thing using the public library.
Still can (libraries still exist plus the internet) but finding the online resources gets more and more difficult; the fact the author mentioned youtube videos instead of someone's old-internet style "all about pvc pipes" website [0] or comic sans plumber sites [1] is already telling.
[0] https://tomtilley.net/projects/pvc/
[1] https://www.plbg.com/
(found via https://wiby.me/)
What do you think it's telling of?
I'm not sure if the sites you linked would have been immediately very helpful in my specific case. You know this is for a pool, right?
The first is called "Everyday Uses for PVC Water Pipe" and has some cool ideas like using PVC for Wiimote holders or tridents, but I don't think that's super relevant for this project.
The second is a great forum which I'm already familiar with, but again, the topic is household plumbing and from a brief skim, none of the topics mention pools. I think mine would have been out of place.
I appreciate you trying to help!
Definitely! My understanding is everything I did has been pretty well established pool maintenance for decades now. I'm sure there are loads of good books on the topic, though I'm a good 45min from the nearest library so it wouldn't have been my first choice.
Fellow pool owner who has always drawn the line at “only pros do the pvc cutting”. I’d love to hear more.
I answered something similar above, but I'll paste it here too:
This was for a largish (40,000L) above-ground pool with no hookup to my home's plumbing system. The water was all pumped in from a water truck. The previous system was also installed by a non-professional and was mostly tubes. It leaked to all hell and looked generally redneck and awful.
First step was draining the pool and doing a nice deep clean. Then I ripped out all the original plumbing until it was just the pool outlets, pump, and filter.
I arranged it all and measured the dimensions. I fed the figures along with a tonne of photos and explanation to GPT. I spent a while talking pros/cons and landed on a design which lined up with what I'd seen on YouTube. I had it prepare me a full shopping list of PVC, tools, cements, etc. all linked to a local pool dealer. I picked it up the next day.
The PVC was all cut with a chop-saw then primed and cemented together. I found this part easier than I would have expected. I put a layer of TigerFlex hose between the PVC manifold and the pump/pool/filter inlets so it had some tolerance.
We've been swimming in it all summer, no issues whatsoever so far.
I feel like if you had wanted to you could have done it before. When I got a pool in 2020,there were still old school phpBB forums out there dedicated to DIY and maintenance. That same summer, my neighbour who in no way is techie or handy redid his plumbing too.
Yeah, absolutely. The YouTube guides have all been very helpful, a lot of them are done by actual paid professionals promoting their pool brand.
The fellow who'd done it originally is my neighbor and he's a retired teacher. There's a lot of ways to learn this stuff - my one takeaway has been that it's much easier than people make it out to be.
This was my take as well. I just don't buy that ChatGPT was some critical prerequisite to doing a pretty standard DIY task.
There are just boatloads of people who say "Oh LLMs are magically good because of democratizing access to information"
Except, for people who were gently motivated, that information was already pretty well democratized by libraries. You could trivially go to the library, and get whatever books were published on a topic, even if the only copy was on the other side of the country. Tons of the famous names from previous decades got their start teaching themselves things from a book in a library. It was very common in the technical churn of the 20th century that a new project at work meant you went to the library and grabbed books on a brand new topic and self-taught. This for example is how some programmers in the 90s developed 3D engines.
There was even a short period of human history where it was common to pay a few thousand dollars for a family encyclopedia. I got my start reading an 80s encyclopedia, focusing on the more technical tomes, before I found Wikipedia. The drive to access and learn information lead to me learning about computers in a time and place where a formal education on the subject was unavailable to me. I owe my career to it.
After the existence of Ebay, a few hundred dollars could populate a shelf with the standard reference books and material for nearly any interest. All it took was a willingness to look for books, buy them, and sit down and read them.
Similarly, the internet did the same since the 90s. Specifically, it allowed for non-physical clubs to supplement the fact that not everyone lived in Silicon Valley and could access those rich clubs on niche topics. But special interest magazines were already providing some of that functionality.
The primary filtering LLMs do is provide new access and ability to people who are far too lazy and unmotivated to do the real work necessary to learn about something without being literally spoon fed.
This helps explain why the primary thing LLMs have done is increase the noise floor of information, and explains differing sentiments. People willing to put in minimal effort to learn things had zero issue learning new information in the previous regime, so aren't that impressed when an LLM regurgitates the wikipedia intro paragraph or summarizes a popular reference. They already read that. They note that the LLMs confidence is often unwarranted, and they get reasonable results because they have foundational understanding of the domain and already know which pitfalls and problems to be concerned about, and how to prompt the LLM to make the right choices.
For people who largely are unwilling to take minimum effort to learn something new, of course LLMs feel magical, because Wikipedia level introductions to topics are magic to people who aren't already seeking them out. Of course, the question is, what in the world was previously stopping you from learning new things?
I think insecurity, over-estimating difficulty / effort, and risk aversion are all to blame. I think these things would improve a lot if people have DIY tasks as part of their upbringing or education, if only to get experience with it.
But the education system - at least for me, 20ish years ago - was very much tiered or broken up into classism: if you were highly intelligent or a good learner you'd go to advanced schools where you'd get higher level math, latin, etc. If you were "dumb" you'd get taught how to do woodworking and masonry. At best I was taught how to use a figure saw, drill press safety measures, and how to patch an inner tube (welcome to the Netherlands, this is very important. Or, was, it's much simpler and cheaper to just buy a new inner tube nowadays).
I've love to know more about how you did this. Was it the underground plumbing or your pump system above ground. This is one of the few areas of the home I have little understanding of.
This was for a largish (40,000L) above-ground pool with no hookup to my home's plumbing system. The water was all pumped in from a water truck.
The previous system was also installed by a non-professional and was mostly tubes. It leaked to all hell and looked generally redneck and awful.
First step was draining the pool and doing a nice deep clean. Then I ripped out all the original plumbing until it was just the pool outlets, pump, and filter.
I arranged it all and measured the dimensions. I fed the figures along with a tonne of photos and explanation to GPT. I spent a while talking pros/cons and landed on a design which lined up with what I'd seen on YouTube. I had it prepare me a full shopping list of PVC, tools, cements, etc. all linked to a local pool dealer. I picked it up the next day.
The PVC was all cut with a chop-saw then primed and cemented together. I found this part easier than I would have expected. I put a layer of TigerFlex hose between the PVC manifold and the pump/pool/filter inlets so it had some tolerance.
We've been swimming in it all summer, no issues whatsoever so far.
I hope what you’ll discover to be wrong needs only $5000 of repairs.
I hope your DIY projects go well and you don't need any repairs at all.
There's plenty of value from LLMs, if you treat them as an advanced search engine and auto complete machine like they are. I've used them quite heavily to research things...but these are done best in the hands of skeptical people.
The problem as usual is the mass of rich people trying to profit off of them, not the technology itself.
Answers to a lot of questions, with breadth, depth and nuance. And assistance in carrying out work on a lot of fields, not just programming.
Do not mistake confidence and verbosity with breadth, depth and nuance
This only matters if the answer is correct.
breadth, depth and nuance
If ever there were three adjectives which do NOT apply to generative AI...
Oh the text definitely has breadth, depth, and nuance ... just not sure how correct it is.
It's really good at writing code.
It’s really good at writing bad code - something that superficially works but is badly architected, abstracted, or wrong in a subtle way.
Only a novice would look at ai code and say “wow this is good”.
You don't have to let it design the architecture. Personally, when I'm using AI: I design the software, and the AI implements it. The classes and functions are typed the way I want them to be typed. AI does a great job.
Yeah if you just let it run wild, it will produce subpar results. But that is not much different from humans tbh.
When you discuss design and architecture first, and write that out in a design doc or something along those lines, it works quite well for the most part.
And 9 out of 10 times when it produces some poor results, just asking "is this really a good approach?" or just stating "This code makes me very sad" it will most of the time do a really good job of analysing why that code is bad and how to improve it.
Analyzing? No. It’s just being sycophantic. Try it on code that’s actually correct - it’ll go all “oh you’re absolutely right” and wreck it anyway.
I was talking about style and architecture. In all these cases the code is working as intended either way. There is no such thing as "correct" style and architecture. There are only tradeoffs. You should know that if you have 31 years of experience.
> There is no such thing as "correct" style and architecture. You should know that if you have 31 years of experience.
And you should know that this is not accurate. But at this point this has devolved into a dick measuring contest, which I refuse to do. Have a good day.
For well-scoped tasks I wouldn't say the code is bad, not brilliant for sure, but definitively good enough.
For a lot of problems "quick and good enough" is all that is required. I've used it a lot for managing my Home Assistant setup. Has saved me countless of hours.
In principle I could've done it myself, but I never would have, the time investment required to learn it wouldn't have been worth the value I get from it.
I’ve been programming professionally for 29 years and with the right constraints (strongly typed language, strict linting, adversarial review styleguide, test coverage, clearly defined spec or requirements) it does indeed regularly write good code. All of the other stuff required is either set-once-and-forget or is also AI, and it does require supervision, but it writes good code the vast majority of the time.
Note that I am talking about frontier models and only the last 6-12 months. Opus was really the breakthrough point for me.
I’ve been programming professionally for 31 years. With all those guardrails it writes functional code - but I would still not call it good. But hey to each their own.
I think of the distinction here is about the domains we're coding within. If your software projects are product-y CRUD apps/sites... well, LLMs will tend to perform decently at that, because it's fundamentally pretty trivial work anyways, and there are so many examples to draw from. All CRUD apps are essentially isomorphic, with just some rules to plug in about data validation, business logic, etc. On the other hand, if your software projects involve more-difficult subject matter, you're going to find it struggles a lot more.
Skill issue. No pun intended.
Something is better than nothing.
A pain plate of rice is bad food. But would you prefer starving too death? Or for others to starve?
AI is like eating cardboard.
AI can make fully functioning computing solutions. That's just a fact and denying it is like denying that a bicycle can roll and only four wheeled vehicles can.
Or saying that a car is a better vehicle than a bicycle. That is probably true, but many times for many people a bicycle is all they can get.
Glad you went for this analogy. It’s like an AI producing a board with a beach umbrella nailed on it and four wheels, and saying “meh it’s good enough for transportation, human car designers are doomed”. It’s functional, right?
If that's your conception of a bicycle.
It’s good at all kinds of stuff that is search adjacent, “find amesent parks with water feature within 100 miles of me that have rv parking nearby”, etc
This will of course never be underminded by aggressive marketing
Rote classification doesn't require AI. You described a SELECT statement that was developed half a century ago
Pseudo SQL
But what we should have is a good map software where you could filter by type and distance (How many amusement parks can be in that circle?) and quickly check if there's a RV parking nearby.The hard job is collecting the data.
Ehhhhh it’s faster at typing code. Writing in a full agentic mode, we’ve only experienced slop.
If typing is the bottleneck you're working in too low-level of a language. Use more abstraction.
Here are a couple examples: https://takes.jamesomalley.co.uk/p/build-gemini-build
> Apparently it’s capable
And apparently is more than adequate to its gulled users.
You know Ava from those ads?
Her and I are… kind of a thing, not like dating dating but, yeah we hang out
I installed Ubuntu on an old laptop that I let Claude Code sysadmin. It makes it really easy to self-host open source stuff, and if there are issues it can fix them too.
Funnily enough, I've had a somewhat mixed-to-hostile response when trying to upstream the vibecoded fixes, so I suspect using an LLM to fix broken open-source software (that human maintainers don't have the time to fix themselves, nor the humility to accept an LLM-authored fix) will become more of a thing going forward too.
Oh, it's also been identifying a bunch of patterns in sales data for my business that has been increasing monthly profit consistently since last November (around $4,000 USD, every month, cumulatively so far with no sign of slowing down - could easily be $10k/mo in increased gains by end of financial year).
It can be very exhausting to receive a lot of PRs. especially when they're LLMs. We have to read what people don't often even read themselves. They're often very wordy. Then it hurts all the more when it's wrong.
I've had negative responses to small, isolated submissions that I've heavily QA'd and semi-positive responses to longer ones. I think it has less to do with the "realpolitik" of the code itself and more to do with the maintainer's viewpoints about whether AI as a whole is a positive or negative thing.
I tend to respond quite well to AI-authored or assisted PRs to my project, but to be fair we maybe only get 3-5 PRs in a good month.
I tend to value person-to-person interaction. I've received a lot of purely automated responses, which shows low value to me as a person, and ultimately the project. I'm not disagreeing that code matters, but we're losing person-to-person discourse and the community that comes with it.
Thinking it through a little bit is probably coming from having bounties on a few issues that might be contributing to my negative experience.
Perhaps the problem you're describing is actually Google's bittersweet solution to the problem of not having unlimited storage space to properly index the whole internet.
The whole text of the internet from day 1 and the index machinery for it can be stored in some number of petabytes. I assume the CIA and the Chinese state are in possession of such stores and it is a minor budget item.
Google and good is of some debate. Appearing on a list of 10 links is not a democratic representation of the web. I'm hoping this reckoning is also an opportunity for a better method of web organization
In original Google, "appearing on a list of 10 links" was peak democracy. You needed people to vote for your website by linking to it. The more votes you get, the higher you are. Sure, like every democracy, it had some unfixable fundamental flaws. But lack of democracy was not one of them.
99 % of the Internet was already slop before AI existed it was just written by hand. At least AI can give you a nuanced overview of medical information whereas before you had the exact same content about medical conditions on hundreds of SEO optimized pages. Collective memory, yeah alright.
Agreed except for the "immense value" part.
Whatever AI is doing to how we find / share information aside, I do think that one under appreciated effect is on how software (and other things) get built. Lowering the barrier to entry for writing code and other tasks feels democratizing , but we may just not have had enough time to see how that hopefully has positive effects a decade out.
Its funny how fast we forget that when Google came out and arguably now many would laugh at the statement "all the good, companies like Google brought to the internet"
I've noticed these big tech companies use the word "democratize" when they do something that looks more like commoditizing. Flooding a market with supply consolidates their own power, by suppressing other economic actors' bargaining power and making quality controls uncompetitive.
Flooding the market with supply while quality might not be the same is a real risk.
> Lowering the barrier to entry for writing code and other tasks feels democratizing
It really isn’t doing any of that, though. When AI gets it wrong, that person needs to understand why it is wrong, and without that upfront knowledge or skill of reasoning, it’s much less direct to actually build something in the correct ways.
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
https://nwu.org/nwu-denounces-cdl/
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
> No. The court specifically determined that the Internet Archive was guilty of unauthorized copying.
You're not wrong, but you're treating “guilty of unauthorized copying” as a statement of physical fact when in reality it just means it falls under an arbitrary rule invented by humans (namely, the law that defines unauthorized copying). This rule is ambiguous at its edges because it's not written as an algorithm or equation. It was perfectly reasonable for Kahle to believe that the rule can be interpreted in a way that it wouldn't apply and, by dragging it through the courts, have that interpretation be made the established one.
Even though the court has now established a competing interpretation, it is still not unreasonable to ask whether the law is fair and just under this interpretation. I feel that it isn't and should be changed.
Seems inevitable, doesn’t it? Expecting otherwise would have been hoping that notorious atheist Richard Dawkins somehow spared one specific god. Making websites accessible with history ignoring copyright is sort of what it does. That he would do it with books seems entirely in keeping with the philosophy.
I was initially confused what Dawkins was doing with books, until I realized that the "he" in your last sentence was Kahle, not Dawkins. Might want to edit your comment to put his name in, because otherwise you have a pronoun referring to a person named in a different comment (rather than the person named in your comment), which could get quite confusing if more people comment on the parent and their comments push yours down the page.
Fair comment but sadly I noticed yours past the typo fix window. Fortunately it has your clarification. Thanks for being able to repair.
> The court specifically determined that the Internet Archive was guilty of unauthorized copying.
That's what the "successfully" in "successfully sued" means.
I agree that IA should have never done CDL but for the opposite reason: They should have never embraced DRM. Either make things available unrestricted and be prepared to defend or don't release it at all. People being unable to loan works during the pandemic may have just been the push we needed to get more people to see the ridiculous onesidedness of todays copyright laws.
And secretly upload all the files to Anna's Archive.
> The court specifically determined that the Internet Archive was guilty
This doesn't seem to be a contradiction. Sometimes the courts are wrong or even the law is wrong. It is just a label.
My sister, a journalist, mentioned to me that she only uses google search because she had learned how to get information typically only Google indexed in the country she lives in, in a way it was not exposed on chat bots. She often has to search for information like Old govt forms released as public record with a fixed a certain format photo scanned into a pdf and indexed by Google were often on the second page of the search and beyond. But they are there. She knew how the forms looked and what bigrans and trigrams matching a certain part of form for a certain piece of information to search for and Google search has it. Like an official order on a tender notice for some government department which is no longer in the .gov.* website gave her the official's name and then she could track down who to contact in an office... ChatGPT and other bots don't have it. Some how all these government documents became part of the government record and are the key for her to do her job.
I sincerely hope google wont stop indexing that stuff just because of a PM in search "de/re-prioritizing" ranking in a way that makes this impossible.
This is a conundrum I always found interesting. If you know how to use a search engine (i.e. knowing how to use operators and structure a search query), you're almost always able to find what you're looking for very quickly, and in most cases (well, before SEO), the results are high-quality. You'll spend the same amount of time trying to fact-check an LLM (since you'll likely skim the articles it used in generating its response _which you would have done anyway if you used the search engine directly_).
I actually took a (required) class in middle school that taught us how to use a library. Amongst other things, the librarian taught us how to use Google effectively. Everything I learned then (this was in the early 2000s) still works today, since the process of using a search engine hasn't changed very much since its inception.
So many people never learned (or never cared about learning) how to use a search engine, thus why we're here today.
Google as a search engine got much worse, especially in the last year or so. The index also got noticeably smaller. My 20+ years of Google Search experience are now failing me completely.
> you're almost always able to find what you're looking for very quickly[...] You'll spend the same amount of time trying to fact-check an LLM
I've been an expert Google user for over a decade and I can only partially agree with the first statement, and not at all with the second. Yes, a search engine alone is fantastic at finding things based on keywords if you know how to invoke it properly. However, there are lots of things one may want to find out which can't be reduced to a keyword search, because you can't have the vocabulary to search for it directly unless you already know the answer. Indirect questions such as "framework options to do x and y in z situation in this language". The best you could hope for pre-LLM was to find forum posts asking the same or a vaguely similar question and comparing a lot of options, finding out you picked a dud after spending an hour on it because it's fundamentally incompatible due to reasons, searching again, etc. It's hard to overstate what a massive improvement LLM's are for this kind of search to find and compare options for exactly what you're asking for given the context of your situation.
I guess it really comes down to how much faith you're willing to put into the answers LLMs are putting in front of you.
If you trust them blindly, then the vast experiment improvements are obvious and apparent.
If you don't, or if you're the kind of person that likes to research your sources, LLMs are a speed bump.
> because you can't have the vocabulary to search for it directly unless you already know the answer. Indirect questions such as "framework options to do x and y in z situation in this language". The best you could hope for pre-LLM was to find forum posts asking the same or a vaguely similar question and comparing a lot of options, finding out you picked a dud after spending an hour on it because it's fundamentally incompatible due to reasons, searching again, etc.
I disagree with this. Stack Overflow (pre-moderation insanity) forums and the like was and is great at finding answers to questions like this. Reddit threads were also useful for this sort of discussion. Much learning was had while reading through the comments on my way to the answer. Sometimes, doing that refined or re-aligned what I was looking for, as is common when doing research.
Again, it comes back to faith in LLMs. Sure, I can ask an LLM to give me a comprehensive overview of web serving frameworks for $LANGUAGE. It's up to the user to determine how valid the information being put in front of them is.
LLMs can also be extremely confidently incorrect. Example: I used an LLM recently in "research" mode to outline how a solution I sell stacks up to the next biggest competitor in pricing structures. It gave me a lot of (too much) information in a readily-digestible format, including, surprisingly, the "agreed-upon" price of the competitor's products per SKU.
Pricing for enterprise sales contracts is very dark arts, so I went to the sources attached to the result to confirm those numbers. Lo and behold, the prices I was given were nowhere to be found in any of those articles.
Meanwhile, I used a search engine manually to see if I could find a leaked price book using "filetype:pdf" operators, mostly for grins. Found it in 15 seconds. That still wasn't applicable to what I was looking for, as it was for an industry different from mine, but it was there.
I could have told the LLM to deep search PDFs (despite telling it to "ultrathink"), but at that point, again, what is the point of using LLMs if I can do the work myself?
Would be interesting to know if those same long tail results come up in Alt-Power [0], which also uses Google's index. So far I get what I ask for, but would be reassuring to know the whole long tail index is indeed shared and I'm not missing relevant results.
[0]: https://altpower.app
hum interesting never heard of this!
Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.
The Internet has been shrinking massively. My earliest experiences with the Internet were discovering the world of hobby OS dev around the turn of the millennium, when I chanced upon someone’s personal website talking about their OS, with source code and screenshots and dedicated forum. My mind was blown. I spent two years finding hundreds of small websites dedicated to the topic, hung out on IRC communities with other teenage OS nerds like me, and of course participated in the nascent osdev.org forum. To note that all of those websites were readily found through Google, and interlinked with their own topic webrings.
Today everything has disappeared or has been conglomerated into siloes, sanitised, focusing on engagement. You have YouTube videos about it (which is more cheap entertainment than actual education), you get some posts here once in a while, there’s Reddit where all intelligent discussion goes to die. IRC is a wasteland of idle bouncers. Then the LLMs arrived to kill what is left.
Who says the Internet is a vibrant place today mistakes flashiness with depth. It’s all empty calories, just makes you hungry for more, never satisfies.
On a whim I watched my favorite childhood movie last week, Hackers. It's goofy in some ways, but man it captures the "wild west" feeling of early and mid 90's internet. It was just you, a slow connection to anywhere, and open ports all over the place. Right after I watched it, I dusted off an old hub, connected a few external usb-to-ethernet adapters to my work PC VM's, and now run them through an OpenBSD packet filter. For no reason at all other than to feel that again: me, watching packets, having total control. Hitting a wall and having to read a manpage.
I don't really have a point I guess, other than even after being steeped in a dead internet for years (with a slow decline spanning at least a decade arguably) I need to approach what I think is the internet in a completely different way. As in, not at all besides what is absolutely required for work. We're ants in a jar now, not cowboys like we used to be.
May I suggest you to look into mesh networks? I am a huge fan of Reticulum. Using it feels like being a pioneer, the scene is very welcoming.
The pitch: it is network-agnostic. The same mesh network runs on the Internet or through LoRa radios or any other physical layer than allows the exchange of data packets. It scales from private networks to global meshes. It's the wild west. People are excited, and eager to grow further.
https://reticulum.network/
I always find these mesh networks landing pages a bit lacking in information, and I wish they weren't.
How is this different to, from example, Yggdrasil Network[0]?
https://yggdrasil-network.github.io/
Makes me think of good old Yggdrasil Linux.
Thank you! I'll definitely look into this.
That's funny because the internet is also network-agnostic. It's xkcd927.
There are retrocommunity sites. DOS, OS/2 (eComStation, ArcaOS), Amiga (Apollo Vampire), Mavericks Forever. Amiga is quite alive community. If I am to write game, I consider this platform. Steam game will be lost in 100.000 games, and Amiga game will be noticed.
There is I2P. They have been disabling scripts for ideological reasons, so there were many websites without scripts. I have used quite exotic Charon web browser from Inferno OS in I2P. That was 15 years ago. Don't know how it's now.
There is RetroNAS and plenty of other software to enrich home network.
Good point on Amiga. It's truly a vibrant community, still. I follow Amiga Bill and Retro Hour etc.
That movie is ridiculous and still great. At least as a nostalgia pump.
And the red box references and pots patching and such were true enough, even if everything else got hilarious hollywood treatment.
If I recall correctly, the film actually had Emmanuel Goldstein of 2600 and Kevin Mitnick as consultants, so despite the hollywood treatment you have moments of accuracy like https://www.youtube.com/watch?v=4U9MI0u2VIE. Probably the only time Compilers: Principles, Techniques and Tools made it to the big screen. At the yearly 2600 conference, they used to always do a big group rollerblade through NYC.
Yeah it's a whole lot better than the usual Hollywood depiction of hacking which brought us gems like creating a GUI in visual basic to trace and IP address. And more importantly it doesn't take itself too seriously which I think makes the ridiculous parts work just like much sci-fi takes liberties with the science part when required for the story.
Cereal Killers name was Emmanuel Goldstein, too.
That movie still gives me wardialing and 2600 meetups and “voice bridging from a Dennys payphone bank at silly hours” memories.
Interesting, yeah I thought there must have been some consultation there. Especially because there's a couple instances where the hackers rely on social engineering over the phone first, and faking the "coin added" noise on the payphone. I guess a hollywood suit could have come up with that, but I doubt it
"Let's echo 23, see what's up"
I, too, look back fondly on the early web. Lately, though, I have been thinking that one reason for its decline wasn't just corporate interests like we often talk about here. A significant number of early bloggers were middle-aged and elderly people. After decades, they simply aged out. The younger generation that replaced them (albeit not so much among OS nerds like yourself) was less likely to use a real computer and keyboard as their interface to the internet, just a smartphone. Hence long-form text died.
While we're here exchanging old man stories, I remember the web before blogs existed! I guess we can organize the internet into these eras, each of which was in some sense harder and more expensive to find available info than in the previous:
1. Pre-web. Internet is mostly about messages sent to individuals or groups. USENET organizes group discussion into browseable topic-oriented hierarchies, IRC does the same but with lists in fragmented networks. If the discussion exists at all, finding it is easy.
2. Early web. Dominated by topic focused websites, early online shops and personal home pages. Search engines suck and face strong competition from manually maintained topic-oriented directories (did anyone else here contribute to DMoz?), content discovery is mutual and webmasters help each other out by joining "web rings". DoubleClick and AdSense start to funnel small amounts of money to creators, but it's enough to offset hosting costs and in many cases can make web hosting effectively free or even yield a small profit. This encourages an explosion of website creation. Discussion moves off USENET onto phpBB forums. Every organization decides it's a cultural imperative to have a presence on the information superhighway. Finding information is easy as long as you can figure out what topic it belongs to.
3. Blogging and centralization era. The internet starts to rebuild itself around people as the primary object, not the topic or category. Directories die because websites can no longer be categorized by content. Web rings die for the same reason. IRC is replaced by instant messengers that are about connecting people with pre-existing friends, not mutual interest groups. Outside of institutional websites that exist to promote the organization, things become hard to find without highly centralized search engines because nobody is putting any effort into organizing or indexing what they write anymore: maybe you get a few tags if you're lucky. Spam, hacking and lack of SSO causes forums to centralize onto Reddit. This is the peak of the search engine era because you are forced to use Google to find anything. The power eventually corrupts the tech firms and they begin political censorship to benefit the left in 2015 [1]. Enormous amounts of information is deliberately made unfindable as part of a large-scale programme of social control.
4. Social media era. All the same problems as blogging except now the bulk of the content goes behind login walls that stop search engines from surfacing them. Video and podcasts start to matter more, both of which are unsearchable by default. Eventually video completely dominates, as few younger people want to read when they could watch instead. Firefox starts to replace IE6, and then Chrome. They bring ad blockers in their wake which starts to choke off ad revenues, so many websites from the web's first era go unmaintained and eventually offline. This is somewhat but not entirely compensated by the falling cost of web hosting. Social media remains because it puts people's faces next to everything, allowing clout farming and viral notoriety that can sometimes be monetized by becoming an influencer. The only part of the web's first era that really survives into this era is Wikipedia and Reddit, which by this time substitute monetary rewards for power tripping by a small group of ideologically driven moderators.
5. AI era. Information is so heavily scattered over so many tiny sourcelets and search engines have become sufficiently useless that full neural integration of knowledge is required, with LLMs issuing massively parallel and complex search engine queries as a backstop.
What can we predict for the AI era? Institutional websites will remain because institutions still have an interest in getting their agenda into LLMs, but visual redesign efforts will largely cease as traffic stats seen by executives show visits completely dominated by AI. There will be lots of conversations of the form, "why redesign our website to look more modern when 99% of traffic is AI which won't care?" Blogs will go the same way as the thematic websites they killed, disappearing as the authors age out. A lot of effort will be put into finding ways to block AI crawlers to create 'human only' spaces, especially by social media firms, but these will fail because AI will just be integrated directly into browsers and become unblockable - and anyway, the incentives to create will be ignored. ChatGPT style text oriented interfaces will last until inferencing capacity catches up, being eventually replaced by voice interaction and on the fly video generation for nearly all users.
Where we go from here is hard to say. Content creation was most pure in the web's first era, where people with knowledge were incentivized to share it with the world by the promise of a bit of fame combined with ad clicks to offset hosting costs. Ad blockers, social media and AI killed that world. You could however bring it back by producing a new platform that isn't like the web, one where AI and search engines are blocked via technological means (e.g. confidential computing). How much anyone would actually enjoy such a web is unclear.
[1] https://arctotherium.substack.com/p/the-closure-of-the-inter...
BBS's should definitely get a mention in the pre-web step
Most hobby communities in my spheres of interest have been subsumed into Discord. It is yet another silo, but it does feel lively. I don't think I'm contradicting you here, I just feel less negatively about it.
I think Discord is one of the worst things to happen to the web. So much useful, interesting content gets walled off and locked away. It's not public, not searchable, not linkable, not indexable.
I'm never going to stumble across an interesting tidbit of information on Discord while browsing the web.
And even when you do have access to a particular server, you'll often struggle to find something posted a while back, even if you know exactly what you're looking for.
I hate Discord with a passion.
Yes, I share your hate. Unfortunately, the open web is apparently too hostile now for any sizeable open community.
I don't think I've ever tried Discord. Is it like IRC? Or more like a forum
Closer to IRC, although they do have forum-like features. Unlike a regular forum, it has no web-facing presence (unless you rig up some kind of bot to mirror it).
Checkout Kagi SmallWeb !
I love Kagi, I am a paid user since day 1, but SmallWeb is a collection of English-speaking tech blogs, it cannot even begin to compete in diversity with the GeoCities era of the web. (to be fair: I've had Kagi employees telling me their dataset is growing larger and more diverse every day, so worth keeping an eye on it)
Sometimes I use Marginalia's "Vintage Web" search for niche topics; most results are dead blogs and old .edu personal websites that someone forgot to delete, still a vanishing minority of anything one could find in 2001.
Well, just as the "internet" supplanted newspapers, magazines (gosh those classic gaming mags), and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
Why should I have to suffer through Fandom raping me with screen-obscuring banners and "PLEASE ALLOW ADS" just to make some sense of fucking Warhammer 40K lore (written by unpaid volunteers anyway)? instead of just asking ChatGPT what the fuck Globriznaroks is/are.
Why should we support shady companies by sitting through their ads on YouTube videos for minute topics instead of just asking AI for the shit I want to know about?
Why should we submit to the whims of 3 mods on a subreddit deciding what thousands should get to see (fuck /r/AskScience) and then getting low-effort answers or outright trolling anyway? instead of just asking AI?
Bury me, I am ready.
Where is the information going to come from though? LLMs aren't primary sources, they consume and regurgitate.
I guess/hope that's where Yann LeCun and their proposed "world models" etc are going to fill in :)
This is better for you as an individual, but worse for society as a whole. As you've noticed, monetization is the weak spot. AI will not escape being ruined by monetization, but it might be harder to notice when it arrives.
Correct, the enshittification of AI would fuck us up more widely and deeply than the enshittification of the internet
but..that's not a problem inherent to the technology or medium itself.
That's a separate problem that's the responsibility of laws and society to solve
Fandom and Reddit are not the old Internet, they’re ad-driven middlemen that create nothing themselves. Self-hosted or at least self-maintained sites were the old Internet. After that point, the rot had already set in then, it just took until LLMs for it to metastasise.
> Fandom and Reddit are not the old Internet
What exactly WAS the "old" internet anyway?
100 half-baked sites hosted on Geocities, Yahoo, about pointless stuff, covered with gif-vomit that looked like epilepsy simulators?
Serious question: What do the rose-tinted glass wearers actually think was of objective substance on the old internet that's nowhere to be found now?
You can find random pointless stuff now too, just that except Geocities/Yahoo it's Intsagram/TikTok/Twitter etc.
If you mean self-hosted websites, they're still here.
If you loved all the Flash toons on Newgrounds etc there's unironically a lot more shorts and animations on YouTube now, if you but search for them (I suggest Weebl, David Firth, Sechi, to start with, and let the algorithm soak up the weirdness)
Popular wikis such as Minecraft and Runescape have successfully migrated away from Fandom. Note that Fandom leaves behind the outdated old wiki and refuses to allow deletion because it drives their revenue. After a few years, the Fandom wiki is severely out of date and loses traffic.
What we actually need is a browser that filters out bullshit.
I really wish Memory Alpha would migrate off of Fandom.
I've been using duckduckgo exclusively for two weeks now, and I haven't needed to go back to Google even once yet.
Duckduckgo has been incredibly worse with results for me for the past year. I had used Duckduckgo for about a decade, and now it is just littered with AI generated content for search results. I recently switched to a search engine that uses Google results.
That's 100% correct. Since Google's "helpful content update" Webmasters get massive amounts of "Crawled, not indexed" reports for anything Google considers "more of the same" or "thin content". If you're not an authority on a subject, simply meaning: you already rank for similar content, or if you don't get links from more popular domains, your content is in the abyss.
It's all under the guise of "We're fighting SPAM", but the algorithm (or model) they use is heavily skewed towards intents (actions) and brands (because they 'trust' big names).
And it's not working.
A simple, short informative blog about a tool you used that could be of interest to max. 100 people on this planet is no longer getting ranked, if it gets indexed at all.
Those posts tick all boxes: no incoming links, no authority, thin content.
It changes somewhat between "Google Updates", but it's pretty clear that it's no longer working.
Multiply the 100 people not finding that post by millions of queries and it's now a big problem for Google.
My dad asked me the other day if Google got worse because smaller companies are not paying Google enough money.
He didn't see any difference between ads and content, because all results are now big brands only.
"Helpful content" is such a misnomer. It removes all helpful content in favour of AI overviews and only shows intent-driven, commercial content.
We're watching the end of Google's hegemony for sure.
This is fascinating. I haven't really kept up with SEO for years, but was recently helping an academic publisher setup their web presence and ran into exactly this issue: Google was/is refusing to index the actual journal articles on the new site and they were moved into this liminal "Crawled, not indexed" state before vanishing entirely. Comically, it's now easier to get your content indexed and served with a visible link by ChatGPT than it is Google...
I can relate to this so much. My interests tend to be pretty niche and I have little interest in most of the major online websites. The scale at which Google has hollowed out the non-corporate internet boggles the mind.
The earlier web era felt like this unimaginable realm of freedom and exploration to me. I stumbled on countless novel sites that were interesting, helpful, and/or entertaining. That content has decreased by orders of magnitude since then. These days most searches return a full page of SEO slop that all summarize (badly) the same source from years before. There's usually zero new information, personal touches, or community attached.
> We're watching the end of Google's hegemony for sure.
Who is standing by to replace them though? OpenAI and Anthropic certainly not, they are burning money in a fire pit to stay alive. There is no way in hell they can afford the compute necessary to replace Google.
If ChatGPT goes bust, they'll find another ChatGPT, not another Google. I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
> I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
Tiktok is cheap to run. AI however, it needs absurd amounts of power, RAM and GPU compute capacity to run... and there's serious constraints everywhere.
Why do people insist on making predictions for the future with today's numbers and constraints? Inference costs are plummeting, mostly driven by Chinese inventions.
Try Marginalia. It's a breath of fresh air. It probably won't find what you're looking for because whatever you're looking for probably doesn't exist in the small web.
About 15 years ago, I won a phone in a contest. Last week I tried to find information about it, but I couldn't. No AI, nor Google could find anything the contest I won it in. When people say "the internet is forever" that can certainly be true, but it isn't for everything.
That reminds me - my wife had a friend who was killed by a shark (no joke). There were newspaper articles, but Google has completely forgotten about this and the newspapers are often behind a paywall, changed their URL's or simply the article vanished from those sites as well, which certainly doesn't help.
I wonder how libraries, who have traditionally been the ones to archive the news, have kept up with everything moving online and now being subscription-holed. Hopefully it's not just the Internet Archive doing this, which has its own problems.
My library system has an aggressive book weeding policy where books that haven't circulated enough in about 2 years get binned. the excuse is other libraries have a copy available for inter library loan, but of course, those libraries have to weed as well.
so no they are not archiving their own core books let alone periodicals
I bet the intelligence agencies are keeping track, though. The likes of CIA and Mossad, at the very least.
Kagi
Kagi is not immune to this (it's the substrate itself that's losing quality, not just search) but it's a decent antidote.
So.. how do we design a Kagi, or a parallel Internet, that is immune to this? There's clearly a market.
The only way to be immune is make ads and tracking impossible or at least not profitable enough for bigcorps to care about putting them there. So, maybe a web without Javascript, images, or videos, lol. Something truly primitive like the Gemini protocol.
But it won't take off (even within our niche communities) without some scriptability and images and videos.
I'll give this some thought. Maybe there is a way.
A lot of the time I use Kagi to find the Reddit post or Wikipedia page I need. There isn't much else.
Kagi can’t solve the fact that the internet itself has gone to shit. The independent forums are largely gone, the mainstream social media’s have all locked down and been spammed with AI slop.
They only repackage what Google et all returns.
Or have they started own indexing?
Here's a Kagi search for the first situation mentioned in the article; you decide: https://kagi.com/search?q=sunset&r=us&sh=CuGH82dPbwxj9_b2i8Y...
They use a mixture of sources. Some HNers like to get very angry because one of the sources is Yandex, which is Russian.
The combination of mass SEO spam/slop and Google's intentional lobotomization of search has lead to it being next to useless. I also strongly suspect Google censors topics at the whims of various government agencies (this was very obvious during COVID and leading up to the 2024 election).
google searches have been terrible for a long time, which is why i have to kind of index all the links and blog posts of use from hackernews....
They're so horrible that I've started defaulting to their AI summaries. And I hate those summaries. It's just that the regular results are so terrible now, and seemingly getting worse at a noticeable pace.
I used to not worry. I was sure that a competitor would come along and fix search. But the longer that's not happening, the more nervous I'm getting that we'll actually lose search. If a few more years pass in the current state, I'm afraid the majority of people will forget what search was like and default to AI summaries.
I've tried alternatives, including Kagi (not actually relevant because there's no way I'm – directly or indirectly – buying Russian products) and Uruky, but they're not good enough.
(Edit: Added "directly or indirectly" about Kagi to point out that I'm not claiming that Kagi itself is Russian.)
Unsurprisingly all search engines seem to be struggling with AI content sites as well. It's rare to get human articles, sometimes rare to even get authorative websites. It's frequently a bot site with a plausible enough name like, potterspainterly.com or medhealthdirect or something with oddly specific articles written in the last year.
This is the root problem. The idea that Google is deliberately sabotaging search seems far less likely than the idea that the internet is mostly garbage and SEO slop.
Google won't index my blog but has no problem indexing AI slop site number 9001. The problem is Google and it's a policy choice.
perhaps this would not be a problem if google did not directly monetize AI Slop... for that matter why even include AI crap in search.
The Google conspiracy has been getting shared for long before AI slop came in to the picture.
The fact is that SEO people got too good at their jobs and filled the search results with junk.
Google can properly classify email spam.
They should be able to fight this as well.
Problem is they are the ones funding the poor quality spam and they're in turn profiting by taking money from advertisers.
Spamming can be found by metadata and reports. There is no "this site is absolute crap" button for them to find useful feedback in searches.
They've had 20+ years to add one.
Google used to penalize websites for SEO hacks. Then they started attending SEO conferences themselves. SEO specialists didn't outsmart Google, Google stopped trying. As the sibling implies this is most likely because they noticed that the easiest way for SEO spam to monetize is ... Google ads.
Brave search works really well, I haven't switched to Google search for months.
I stopped paying attention to Brave years ago when it started fiddling with advertising-linked cryptocurrency and content injection. Is there any reason I should extend it any trust now?
Thank you for your suggestion. I think I've discounted Brave automatically because my brain is numb to the dime-a-dozen chromium browsers out there. I'll definitely give the search a try!
If I recall correctly Brave scrapes the web via their users, cannot be individually disallowed in robots.txt and Brandon Eich is conversing in a pretty hostile manner in every thread about him or his company.
I'm not in the brave ecosystem otherwise nor a big fan. But the search engine was competitive with google when it was still cliqz, before it was shut down there and the leftovers bought by brave. And it still works really well.
Even if brave were problematic it would be the lesser evil to me.
This is just papering over it individually, but https://github.com/iorate/ublacklist is the only thing that makes searching bearable for me now.
Probably what is happening here is that in the race for AI, which, whether we like it or not, means power and control, Google crafted things in a way that it does not do "self-competition" by their old search engine. Idk, just throwing ideas aloud here.
Kagi is not russian. The founder is from Serbia and the company is american.
PS: I hope not many people are the type to see a slavic name and conclude Russia.
Kagi pays Yandex for data. Yandex is definitely Russian.
Yandex also moved outside of Russia and there only one vendor.
You need to get over your Russophobia, because Kagi is actually good. Or else make your own Yandex.
How is Kagi indirectly Russian?
They are repackaging bunch of search engine results, including Yandex (and paying for it).
Yandex is Russian.
I would not describe it as Kagi being indirectly Russian.
> I would not describe it as Kagi being indirectly Russian.
What do you mean? That Kagi is directly Russian?
Paying a russian provider isn't against the law and at some point we need to bring Russia back into the warmth of the West.
Perhaps “at some point” is not “while they’re still waging a war of aggression”, though.
That's certainly the goal, to drive more users to AI by making the other option worse. The real question is when do you put a stop to it. When brain implants are 2x as productive are you going to say nah? It's times like these were you are supposed to take a step back and consider what you are actually producing and why.
> there's no way I'm – directly or indirectly – buying Russian products
Why not?
In my case: Russian state is my direct enemy and most likely to invade my country. And would do it if they would consider success likely.
Previous wars with Russia were obnoxious with very bad consequences, so I dislike idea of even very indirectly funding them.
And I support actions that are harmful to Russian economy, also when they are harmful to me - as long as it is not too badly balanced. As this is much cheaper than directly participating in war.
(I am from Poland)
PS
Yes, I understand that at some point there are some indirect effects that you cannot avoid.
I also understand if for some people paying Kagi that pays tiny fraction of that to Yandex that is paying taxes in Russia which funds their wars is too tenuous connection to care.
I think at this point Russia is failing so hard they don't need a 100% boycott any more. 99.9% is enough.
>Russian state is my direct enemy and most likely to invade my country. And would do it if they would consider success likely.
What makes you think so?
>Previous wars with Russia were obnoxious with very bad consequences
What do you mean?
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
Maybe the curl thing is because they are happy to let you do some light scraping. What they want to avoid is bots directly crawling the page interactively. No one seems to be blocking chatGPT when I promot it to use it's web search skill anyway.
Cloudflare specifically has a block for LLM and AI training bots now.
Not sure of the effectiveness but it's there.
Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
I think the other thing they're trying to do is get most of the internet to send them all of their cleartext traffic. Expect in 2040 the PRISM2 docs will get leaked by some Eduardo Rainedon and we'll find out Cloudflare was the NSA all along.
While they themselves announced an AI bot, the irony is palpable. In reality they just want to control who gets access to what, themselves excepted.
It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.
The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up...
Continuing on your suggestion.
There could be open source tooling to create custom private "closednets", with
- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned
- the rules of the closednet
- search engine with opt-in scraping
- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.
etc.
The first closednet could be Hacker News.
I think you'd struggle to keep LLM bots off the network unfortunately.
If it had any real value, anyway.
Small, truly private communities could be an interesting thing though.
How I can strugle to keep them off? To peer to network, you need to talk to human. Arrange L2 connection, assign IPs (only static). If its leaf node, we are done. If its another network, we need to form BGP connections to exchange routing.
Yeah, Network by Humans for Humans. Thats why Im not interested in all those IoT/Auto networks when you just connect and stuff automagically configure. It looks nice at first glance, but you loose control. F2F works way better in that matter, like RetroShare, but I never investigated it much.
It doesn't have to be an IP-layer network. A website that you need to log in to view works just as well.
Nah, it needs to be IP. IP is well estabilished protocol, everything speak it. Once you set it up, you can use it whatever you like. Web pages, gaming service, IRC, Mail, P2P confereces, everything. Everyone will bring it own slice to the pie. You love networking, became PoP and peer and provide access. You just want content? Connect to closest PoP, get IP + DNSproxy and vioala.
The networking part exists, it's called DN42. But it's hard to set up. I don't see why you'd use this instead of an application layer password.
I know DN42, but this network is more oriented toward R&D and experimenting. Yeah, its not for your avarage Joe. But your avarage Joe can buy connection from VPN provider and use it, and so we can provide user friendly PoPs with minimal skills needed to setup, supporting different VPN software.
I just used a VPN yesterday and the NY times blocked me because they think I look like a bot. It gave a couple possible reasons, one being "a bot was also using this IP address".
VPNs are great for torrenting but any serious website like an online bank or web email provider will turn you away. They claim it's for bots but really it because they only want customers they can track.
Me, I'm just scraping the parts of the internet I like, toying with local LLMs… ready really to just shove off.
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
> are those that are only hosting content for the ad views.
I might never blog/publish code again amidst all this. I never had ads on my sites. I am not alone in this.
Those hosting content for ad views, or for fame or other personal gain.
Sites like Wikipedia or developer documentation pages which exist to distribute knowledge for its own sake don't have any reason to care whether that knowledge is consumed by a human or a computer being used by a human.
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
Are you actually saying you'd be OK with that?
Yes, that would be fine. I write primarily communicate ideas, not for credit or fame.
Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866
Sure they know about you if you ask, but generally they won't credit you if they cite an idea from their latent space that came from you.
That's fine! Humans who read my blog typically won't credit me if I help inspire them either.
(I was pointing out that "they'll train on a version that strips out you as the author" seems to be is incorrect about how training works)
It knows you, but it doesn't know me. Perhaps you are legitimately noteworthy enough to not worry about this!
My experience being on the searching end is that these things are terrible about attribution of where they find anything. Which has bad consequences not just for authorship, but for correctness (which is the usual reason I'm poking at them -- they're being wrong again). This makes a lot of sense when you consider the massive, massive compression that's got to occur during training, but it's still frustrating.
Not the OP, but I suspect no one goes to my website anyway. (I write nonetheless.)
There's massive differences between "my site isn't hugely popular, but I get some readers", "my site is up and findable but genuinely no one visits except scraping bots" and "it's available, people would visit if they knew it existed, but they're not being offered it, they're being offered bot distillations with no reference back".
The last one is the worry.
The middle is... where we all start.
The first is not a bad place to be, all things considered!
I do! You invented the crayon picker in MacOS..?
So I'll write a lot of falsehoods to poison the AI. Like the urban legend where the sky is blue. I'll say the sky is blue, AI will think it's true, and regurgitate it to unsuspecting users who will see that AI is completely unreliable.
Your level of thinking is defined by ICD-10.
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
GENIUS
Thanatic drive masquerading as transhumanist virtue signalling.
I mainly shared my projects for learning, discussion and bragging rights.
LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention.
Worse, there are some PRs that seem fully generated ...
So i mostly stopped sharing and started pulling my old repos offline.
At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper.
There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
My parent wrote "what incentive is there to publish anything new?" and I described my motivation, but of course people vary a lot in what drives them.
> "humans would visit the website and the creator would get some reward"
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves.
If "getting the knowledge out there" is the motivation then it shouldn't matter whether the knowledge gets filtered through an AI or not.
And yet, inexplicably it does.
> it's just about the getting the knowledge out there that's motivating the creator
In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever.
AI will mutate the knowledge and eventually produce some mangled version of it, mixed in with ramblings from a random reddit post and half a paragraph from a copyrighted book that its fascist creators stole.
"the AI will preserve the knowledge forever"
no, only some mangled form of it
> Your thinking
Should be "You're thinking".
Maybe I'm just making silly misspellings like this on purpose, to prevent someone from fingerprinting me based on my writing style. :-)
> Should be "You're thinking".
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
> This should be "This should be 'You're thinking'." don't you think?
Reflexively, I think it should be more like ...
... but then that's just me, in [my] quirks mode.Did the operators of HN mention anything about a reward structure for posting comments here? I'm sure you can see how that's still attractive to some.
I finally got the downvote powers yesterday, and I was irrationally pleased about that.
It's not even about votes, I don't even need votes. Whenever I write something that I am happy with, I read and reread it imagining I am reading it as a third person. Sometimes it forces me to rework my arguments. I wouldn't write to convey my ideas through a chatbot. And that's what this post is about-- killing the internet and with it decimating any audience you might have accrued if you had something to say and you published it on a website.
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
At least LLMs almost always transform the original - it usually isn’t as straightforward as “Here’s the original but without the ads that pay for it”.
But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.
Gemini can just consume the device documents. There's an incentive for device makers to publish this content.
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
That is next earnings quarters problem is the approach being taken
> incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.
Well companies are starting to put hidden ads in text content if the user agent is an AI crawler
People write and create regardless of profit motive, it has been that way for thousands of years.
People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired.
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
I had a flippant answer which was “who ever wanted to write for a machine in the past thousands of years”
But maybe that is the future.
Every country starts erecting their own towers of babel that we talk at, and it constantly compresses our conversations down to the most effective distribution of weights.
At some point talking at the machine becomes a high status job, and we give respect to the people who whisper to it the most.
They don't write for machines, they write for themselves.
Why write publicly then?
Theres many people I know who write notes that I know would be great to read. However they never publish them.
So… are we saying that the only public writing in the future is meant to be consumed by the machine?
Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
Actually it grew out of censorship and monopolies: https://en.wikipedia.org/wiki/Statute_of_Anne#Background
Cheap access came from the invention of cheap printing . The laws were passed to restrict it.
No, cheap books came from the invention of cheap printing. But the quality of works was going down because printers were just printing with zero copyright protections. Copyright is what brought the wealth of works worth reading that were then printed using cheap printing. If the only money is in private works for private libraries, that is where the quality stuff is going to go.
Making the avenue of creating for the average person also the avenue for the most income was huge in creating our modern literature landscape.
> Copyright is what brought the wealth of works worth reading that were then printed using cheap printing
Evidence for this statement?
> If the only money is in private works for private libraries, that is where the quality stuff is going to go.
Evidence that this ever happened?
> Making the avenue of creating for the average person also the avenue for the most income was huge in creating our modern literature landscape.
That only leads to higher quality (as you claim) if your definition of higher quality is "what the average person buys".
It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there.
Disagree. The printing machine was killing the industry allowing copies where people earned nothing. Copyright was created to ensure quality works existed to be copied.
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
> Funny, I published information in the hopes that humans would benefit from it.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
That works now because there are human made sources that the AI can find and summarize for you. But now there are no incentives at all for humans to write anything on the internet and if they do the content will be buried by hallucinated content someone else posted at a larger scale.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
Isn't this only a transitionary problem though? Right now, during the transition there is no incentive for humans to write anything, it will get drowned in AI slop.
As time goes on, more and more people will recognize this problem and we'll develop new ways of measuring information quality and trustworthiness. Nothing about this problem is fundamental, it's just that we're in the middle of a very chaotic transition.
Yes. When my kids wonder why I don't hate AI the way they do I tell them it's because I hate Google even more.
To be more precise, I hate the SEO shithole the internet has become, that Google serves up, that Google facilitated, indirectly created.
(I really don't have any tears to shed if there is a death of the Corporate Internet™.)
AI seems to be on the same trajectory? Search was very useful in the start also, until it became entrenched. Then search placement became a target, and they are just focusing on extracting rents. All way paying the content providers zero or near-zero. And with years of that dynamic, we end up where we are now. It was the same with "social media". The same will happen with AI. AI is a power for more enshittification - being currently less shit than Google is (mosy likely) temporary.
It's definitely temporary. Remember there are no ad blockers for LLMs.
You're right, this is the best it will ever be. But there likely will be ad blockers for them -- local LLMs that filter for any brand placement, etc.
They're probably going to figure out how to aggressively monetize and enshittify LLMs at some point.
We're likely at the "golden age" of LLM-assisted web searching and summarization.
Hopefully open models keep it cracked open, but expecting enshittification is always the safe bet these days.
For the free to use LLMs, for sure. But if I pay $20 monthly, why should they enshittify it?
$20 a month is not nearly enough. The current finances require AI companies to make far more money than that to not implode.
look at every paid streaming service...
To be fair those enshittified because of copyright owners, not because of paid streaming services.
Because that’s what always happens? Because their raison d’etre is to squeeze you dry?
The amount of money you pay is irrelevant if you don't leave the platform once they starting adding revenue streams that degrade your experience.
Is there some search engine that you think could have become popular and not ended up with SEO optimization?
Only if ads were not a thing.
One you pay for yourself !
Very happy with Kagi personally
>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I don’t agree with the parent that the solution for SEO spam is paying for the search engine (I think there are other good reasons to do this though), especially since people have been doing things like naming their company “AAA Auto Repair” to be first in the phone book since before computers ever existed. But the person you originally replied to does have a point in blaming Google for the problem. Most SEO spam sites make their money from ads, and Google are the ones who run the ad network, which means Google are the ones funding them and creating an incentive for them to exist.
It means your search engine is incentivized to fix it.
It's nice to be the customer instead of being the product, for once.
> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
Ah, the argument is that Google intentionally avoids showing you the most relevant results, or at least avoids "solving" the problem of webmasters who attempt to 'game' the system.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
> I see no reason it would have that effect.
Agreed.
Google did not really lost to optimization, they enshittified to make you spwnd more time on search (and see more ads)
Ah, the argument is that Google intentionally avoids showing you the best results, so that you spend more time searching (and being exposed to ads)?
I'd frame it more that Google is incentivised to show you the results which are most lucrative for them to display, rather than the ones which are most beneficial for you to see.
Incentives. Which is why platforms (such as search) should never be allowed to be in the same company with things built on top of them (such as ads), if the combined company is significant for an important market.
We used to know better, Standard Oil vertical integration was dismantled.
That only works when there's an abundance of documentation for your specific device. The moment you're on a more recent version of something and have a weird issue, all hope is lots. You get stuck in loops because all LLMs keep recycling old advice that no longer applies. This problem will only get worse and worse as people no longer as questions on public forums, so answers are not publicly available either.
While I understand that you feel good as you got the device configured, you would have been better off without gemini. The 'old' way as you call it would have resulted in you knowing the backgrounds and the inner workings of your router, making maintaining it a breeze and helps you actually understand your setup. Besides that, it would probably trigger you to rethink some of the things you now blindly have implemented because gemini did not show alternatives nor reasoning behind it. (something that most definitely would have been documented on the source pages)
I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.
> I don't know kubernetes and never will.
He said, proud of his own ignorance.
Proud? No, but I am not ashamed of it also.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
What on earth LLM are you using and do you tell it which distro/version of Linux they are supposed to be working with + give them access to web search or man pages? I haven't had this issue in like a year and a half.
By default I use Leo, and I do specify the distro. I tried to qualify by version, but when I do that I lose out on a lot of correct information for things that haven't changed in awhile.
Leo is not even the LLM, its just a browser extension typically running a very weak/cheap model. Why not use claude code so it knows your system without relying on you to give it accurate info? I guarantee your results will be night and day
I've had the opposite experience giving claude code SSH/ADB access to my devices. Not sure if it's the model itself or the harness, but I haven't had to manually do a sysadmin task in months.
Have a look at https://github.com/ThorOdinson246/whatisit-nl2sh
It seems to get shell scripts right most of the time.
Yep - it's awesome, the problem is that Google isn't sharing the revenue with the context creators anymore - over time, unless fixed, this will decimate the knowledge base it feeds on.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
Yeah I mooch gemini off of google but they've def made a few cents from me via ads. Also apparently I'm a rarity because I don't run ad blocking (except firefox itself does a bit according to a few websites). Its a very good deal for me though.
I hate to break it to you, but we are at the start of enshittification circle. Once Gemini has monopoly you will reminisce with joy google search results.
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
Believe me they're making money off you one way or the other
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.
YMMV.
Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
also great for Windows-Log Exceptions to decipher the mostly cryptic error messages :-)
You deserve what's coming for you.
Thats probably the kicker. I think about this some, also being myself a gemini chat mooch. How good will gemini be in a years time? Its great now but will there be an incentive at some point for it to become shit? Free tier is awesome now, but is it a loss leader? Training on my chats, maybe in a year my chats won't be worth the free access?
I built a simple Claude Code container on my homelab to do the same. I can just SSH in and get tech support when I need it.
AI is on the whole bad but it’s impossible to forget that the ostensibly human-curated Web as presented by modern search engines is bad too. Invasive ads, popups, and worst of all the substantive content has a nine inch frame of SEO filler and a three inch picture (what you were after). And some things are not even high-tech slop stolen. It is just old-school verbatim copied from another website and repackaged with another frame.
But yeah, text remix machines are not a long-term solution to that problem.
Imagine someone making an indie router and sell it too you then the AI not be able to surface their doc
no, no, no. you are supposed to romanticize the hunt for correct information /s
Even with the /s I think you misunderstand. If you just want to configure your router, then AI is neat. If you want to learn and understand what's going on in the router as you set it up, then being given the answers basically teaches you nothing.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
Why would most people want to learn how a router works? It's the end goal that the vast majority of users want - i don't want to learn about routing tables, I just want to expose a port for plex, etc. Also, I could have AI craft a fascinating and engaging way of teaching router setup that actually helps me learn instead of wasting time searching through random forums form google results
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
Are they profiting from these subscriptions yet?
I've been using Gemini free tier exclusively for all my AI "needs" and would not be willing to pay for it
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
Come now, copilot is worse in every way
Guessing this is the relevant part: "worst AI model I ever used"
Copilot can use any model like Sol or Opus and is just a harness so not sure why people say this.
Personally it's because i used it when it was brand new and work paid for it and i had no idea what model it was using, my boss just turned it on for automatic PR summaries and code reviews and it was universally dogshit, and the auto complete in my IDE was awful as well.
Well it's much better now.
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
Interesting, I've seen much better results on DDG. Most recently was the search: `site:feeds.bbci.co.uk inurl:rss.xml` which works on DDG but gives zero results on Google. As far as I can tell, Google just decided not to index these.
Yeah I mostly still use Google out of habit but there have been a few times where Google has decided something isn't worth indexing (too niche, doesn't use SSL).
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
I find DDG is struggling to, or chosen not to, filter or derank obvious AI generated content farm sites. Of which there are an insane amount of already.
We thought blogspam was bad, at least it was easy to ignore. It's hard to find authoritative sources for a number of topics, worryingly health advice is one of them.
Idk, Google seems worse there too. I tried "my left hip is hurting".
Google shows an AI overview and "people also ask" with zero search results above the fold. If I page down I see a single search result for clevelandclinic.org, followed by youtube videos and image search results. The next page has a single search result from rush.edu and then "discussions and forums" which has Mayo Clinic and Quora.
DDG also starts with the AI overview (although I disable that) and has two results from webmd.com with deep links to multiple pages on the site, all above the fold. Then the same clevelandclinic.org result as Google but again, adding deep links to other related pages.
I can't comment on the quality of webmd, clevelandclinic, or rush, but Google pushing the user to youtube and quora for medical advise seems worrying.
The problem and what the article is pointing out is that original content is slowly and progressively being replaced by AI content. And since AI gets trained on this content as well it will eventually train itself on previous gen. content that was also AI generated. It is slowly eating the web. Eventually you won't even be able to evade it because it'll be everywhere. Before AI became this expressive I could at least expect someone writing articles, FAQs, blog posts to have some backbone. Now I frequently run into content that obviously was never even checked by a human.
https://noai.duckduckgo.com is a thing fyi
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
> https://noai.duckduckgo.com is a thing fyi
or you can press the gear button -> "Ai features: Manage" -> Search assist
That was true until a few months ago. Now, it almost never has relevant results, and I've given up on it.
It seemed like Google was good ten+ years ago and then gradually trended towards being rotten around 2020-2022. In those days blogspam was king. I want a cooking recipe and it would show me a life story. Or I wanted a OG web game but a link farm would come up top. I think there are leaked internal comms where they discuss nerfing search results to pump engagement and ads impressions. No doubt it would boost ad bids too if businesses couldn't be found organically.
Around that time Bing and DDG were actually better. Then LLM's came along and they started to take things seriously again. Maybe they think the OpenAI threat has abated enough to begin enshitification cycle 2.0.
I know Google started deteriorating. It was in 2008 or 2009. Up until then Google would return 0 results if it couldn't find a document with all the words you specified.
Then it started serving synonyms, attempted to correct spelling, and so forth. Instead of serving up that there was 0 results, it attempted to be "helpful".
For a while, you could enable "verbatim" search, but even that has gotten corrupted.
Their search quality has deteriorated ever since that change. It is sad.
I've used DDG for years (thousands of searches) and when I switch back to Google thinking I might be missing something... I'm always let down. Seriously, the results are pathetically bad now and have been for years.
that control (so I can turn it off completely) is why I picked DDG to replace Google, who force feeds us the hallucinations
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
Yeah, there are times where DDG has like, literally three results. Yet, I know for an absolute fact, there are hundreds of pages on the web that contain the terms I specified. Web search is becoming utter garbage.
Try brave search.
Both DDG and Brave Search are front-ends for Bing.
Brave Search operates a fully independent index (i.e. we have zero reliance on Bing or any other third-party index).
Disclaimer: I work at Brave
I don't think it's really an "AI problem", we just got to the "worse" part of the "Worse is better".
Back in the day one of competitors of the World Wide Web was Project Xanadu. Project Xanadu was supposed to address the concerns like content persistence and version management within the core design. As such it was much more complex, opinionated and centralized.
WWW on the other hand comes with no guarantees - you might get a document in response to a HTTP request, and that's it. But WWW service can be rolled out in a completely permissionless way, and is quite simple - effectively, the contents of the file system can be shared with the world, so e.g. a document can be published just by putting its file into a particular directory within the file system.
Thus Web could get to a "good enough" state much faster and quickly spread all over the world. But its permissionlessness and simplicity lead to downsides: impersistence and chaos of broken links, web search provided by mega-corporations, etc.
WWW evolution was, unfortunately, not "incentive compatible" with features like advanced persistence and identification clarity: there was much more focus on entertainment content and ads
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
> not even Google can afford to devote enough compute to each query to reliably generate quality results
https://www.dw.com/en/german-court-holds-google-liable-for-f...
I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
Otherwise they'd be slurping in their own slop
Reddit has already begun the effort to start poising the well - https://www.reddit.com/r/poisonai/
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
Think of the more malicious use-cases though: The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.
This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.
I think OP is more concerned with we are going to ensure that new data added to the "trusted corpus" is going to be free of LLM taint.
Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly
That’s what is going on right now.
As I recall there are data labelling jobs now for people who have experience working at McKinsey.
> There will come a day (and probably soon)
That day has already arrived, it is already happening.
Gemini has been a hilarious companion to my while I fixed the balance shaft chain guides in my old Mitsubishi triton (mighty max for US readers).
First it told me I could just remove said balance shaft chain as an emergency repair. Sorry Gemini, it also drives the oil pump.
Then it told me I could remove the water contaminated oil caused by removing the timing case by filling the crankcase with hot, soapy water and running the engine. Lord no.
Then it gave the wrong instructions for putting new gears on the balance shafts which meant the chain guides didn’t align with the chain. I’ll do it my way thanks Gemini.
The rest of the mistakes are too trivial to recount and sure it’s a pretty obscure subject but if I trusted it with a topic I’m not familiar with there is a huge potential for damage if you blindly follow it’s overconfidence. I miss normal searching.
Meanwhile, ChatGPT correctly diagnosed what was wrong with my plant from a single photo, identified which leaves I should cut, and annotated the picture showing where to cut and what not to touch.
I honestly expected a made-up useless generated image that matched the idea but not the actual thing.
Guess I’m still living in 2024.
What makes you think the chatbot's diagnosis is correct?
> While the web has always been organized around intermediaries that shape what survives online and who sees it,
This statement, from the sixth paragraph of the article, is something that I would have liked to see addressed more in the article. The article implies that this is something that must always be true, or cannot be changed, and simply focuses on how we could have better/better funded/better protected intermediaries (AKA gatekeepers), and doesn't discuss the possibility of an internet (or part of the internet) without gatekeepers (and doesn't ask if it has ever existed/does exist/should exist)
AI has killed reading-anything-written-after-AI for me. Due to this effect it is probably the worst invention in human history or pre-history.
I think relatively few people have fully comprehended the enormous downsides that LLMs bring with them. I do think they are useful tools in many situations, but the fact alone that they have "flipped" the "takes time and energy to write something useful so others can more quickly read and comprehend a complex idea" equation is very, very bad.
Not just reading, videos have been ruined in many ways as well. My base reaction to seeing anything surprising on video form has switched from "Interesting!" to "Probably fake". Novel events caught on video are now immediately suspected as being AI and discounted by many people.
Same for software for that matter. GenAI has killed my interest in releasing web apps outside of work. Before I used to find value in putting stuff on the web for others to find and enjoy, and it sparked some interesting interaction with other people. Now anything I publish will mostly be consumed by bots and thrown into an AI blender that completely divorces it from the creator.
Indeed, I also don't listen to music after 2024. I'm also not interested in anything my friends have made (with claude or others). Or how so much of creativity is now a prompt away.
A lot of our culture is disappearing before our eyes and people are defending it.
Wait I’m confused- is this because you believe all music produced post 2024 is ai generated/ai assisted?
I think metal will be fine. Maybe it'll become the dominant musical genre in the US, finally?
You can find bands you like and follow what they put out, just like before.
I don't know why you'd ignore the output of talented musicians because AI exists.
You can't complain about culture disappearing if you're choosing to ignore it.
And make sure "follow" means "buy their stuff on Bandcamp or directly from them". Streaming platforms send most of the money you pay to bands you probably never listen to.
You're comparing almost three millennia of text, to a couple of years worth of text. Let's see this play out. Plato also threw shade on writing. Which is an invention that turned out alright for humanity, imho. Maybe it'll turn out that we can push through to a yet-unknown-but-better state, as we've so far done, rather than relying on just going back to a previous state that we were content with.
Sorry for the somewhat sarcastic tone. But the hyperbolism deserved it.
Plato did not throw shade on writing. This often comes up! Where are people getting this misinformation? The Phaedrus is not that long.. just read it!
Also it doesn't even make sense... How could we know this without Plato writing it down??
Nah, AI writing is overwhelming real writing but contributes nothing new.
You don’t seem to have read the comment you are replying to.
The comment they’re replying to didn’t add anything new - it’s just the same old regurgitated ideas and has a lot of misconceptions. Almost exactly like LLM content. Taking the time to respond in detail would be a waste.
I think they understood it pretty well. What in it was worth any other reaction?
edit: I might add that anna's archive and sci-hub are good alternatives to the slopweb.
Nothing they said demonstrated any understanding of what they replied to. They simply restated their opinion with no reference to any of the points made in the comment they were ostensibly replying to.
Am I really the only one who doesn't care whether or not prose was written by AI? Assuming that a text is signed off on by a person, then surely what counts is its substance, not who or what actually strung the words together. I'm constantly surprised by this obsessive need to falsify what is unfalsifiable. And among nerds of all people. It feels religious, like a modern form of heresy-purging.
I care because in most (not all) cases, (a) AI UGC is one- to few-shotted with no concern for accuracy or readability, and (b) why would I read the results of someone else's prompt if I can achieve more or less the same result for free on ChatGPT?
For example, I was searching for a guide for realigning the v-brakes on my wife's hybrid bike and found this: https://volatacycles.com/how-to-adjust-bike-brakes-rubbing/
Now, this is for disc brakes, so not applicable in my situation. I've never ridden a bike with disc brakes, so I can't vet for the accuracy of this guide. Maybe it's totally right; I'm sure someone reading this will know.
However, just look at this article! Zero pictures for a process that _really_ needs it. It just drones on. Imagine following this guide only to discover that the 12-15 Nm rotor torque the article is asking for is actually too high!
A similar article from before AI would at least have (stolen) pictures describing the process. It would also very likely be shorter. Which creates the rub: if I know how to write something engaging and correcting AI's copy will take _just as much time_ as writing it myself, why would I use AI to do it?
In short, AI content is lazy content. My time is valuable and finite; if I'm reading your stuff, I want to know that you at least tried to give a shit about my time as a reader while producing it.
> However, just look at this article! Zero pictures for a process that _really_ needs it. It just drones on.
Buckle up, because the next phase of this AI content generation nightmare will be articles like this featuring AI generated images of a bike repair in progress, except the bike will only kinda-sorta be like the real bike the article is describing.
Sure, I get that argument. I just don't buy it. What counts is whether the article helps you fix your brakes. If it doesn't, it seems immaterial to me whether or not the bylined author sweated over it for hours. Not least because I have no means of knowing.
But the OP already addressed this-he can just ask ChatGPT directly how to fix the brakes. He’s gone looking for a human expert because (for whatever reason) he doesn’t want that.
For me the issue is very clear: I happily use AI to answer tons of questions each day, but if I’m reading your website/blog/article/Jira ticket, I expect real human input.
It should matter whether the author “sweated over it for hours”, because that means they actually considered the best way to communicate something. If they didn’t do that, then they clearly don’t care enough about about the quality of their output and I shouldn’t spend any time at all on it. In fact it’s worse than that because if the author doesn’t respect the reader enough to craft their output, then I as a reader have no respect for their work, and I resent that they’ve taken my attention for the time it takes me to realize it’s ai generated. If you’re publishing Ai content, then I can only presume you have an ulterior motive than plain communication- in the case of OP I’d assume this is for ad revenue or clicks.
There are many types of writing, and I struggle to think of many where a statistical model could feasibly produce an acceptable output for the given intention.
Eg. A poet chooses their words extremely carefully; a good instruction manual is written by the designers of the product (not just guessing at a common or plausible method); a news article has an angle/story beyond just the facts.
If you or OP suspect (suspect) that a blogger's taking shortcuts, and that's a problem for you or OP, then you or OP can just skip that blog and find another that appears (appears) to meet some criteria for labor input. To the extent that this is all unfalsifiable, it just should not matter.
I'll go further. Even if it were falsifiable, why should it matter? An example: I subscribe to a reputable US publication and I enjoy its journalism. If it turned out that one of its writers had (somehow) used AI to generate their article (which I enjoyed) in a click, here's what I would say to them: Hats off to you! How did you do it?!
I don't care if the article took them two minutes or if they used AI any more than I care if they used a spellchecker or wrote it while standing upside down. Why should I? The author put their name to the article and I got something out of it. That is all I was ever looking for.
GP described multiple reasons it didn't help them fix their brakes.
So blame whoever signed off on it. AI or not-AI is their problem.
If the article doesn't help you fix your brakes, it has confused you and wasted your time. Not immaterial.
> Assuming that a text is signed off on by a person, then surely what counts is its substance
Unfortunately that's a big assumption and for many their "signing off" process will amount to "skimmed it and seems ok?".
When I get an LLM-generated doc or runbook, my first thought is that its very possible that I'm the first person who has ever read this. It used to be that writing something took a big time and energy investment up front by one person so that many others could comprehend the ideas with less time and energy needed. LLMs flip this equation around which is a very bad thing.
The problem is that AI writing uses a pretty distinctive style and when you read a whole newspaper that was written or edited by AI it is very monotonous. Also, I can't really trust it because I know that the author probably phoned in his fact-checking.
Because it sucks! It's badly written, too long, and often pointless. I don't like reading tacky, useless garbage
Yeah, I keep feeling the same thing once I know I'm reading something or watching something purely AI-generated - that there are good points being made, but the delivery is just so off that it kind of ruins the points themselves.
If people used AI to get the scaffolding of their article going and then re-humanized the entire article, it might be really good, still at a fraction of the time it would have taken to write it themselves, but no one seems to want to expend enough effort to go the extra mile.
With video it's pretty well hopeless, unless the AI helps accelerate fully human traditional production techniques.
With music, there might be a middle ground, I have had some success getting AI to generate good sounding loops to use in electronic music production, but it was just a matter of brute forcing enough outputs to finally get something decent, which isn't fun at all.
Why shouldn't people be upset if authors are being replaced by like a handful of the same ghostwriters?
An LLM can't convey your ideas and your voice better than you can convey them to the LLM, right? It's a middleman between you and those you want to reach that adds more points of failure, an additional entity that must be understood and made to understand, or else meaning is lost.
EDIT: Never mind the motivations behind a majority of LLM-generated articles, which is to be One Unit Whole Content, a colloid for ads.
I would be fine with it conceptually, if the result was not as grating as Claude's ramblings about load-bearing seams and the real shape of the problem.
It's just exhausting to read, and it displays a lack of effort put into communication.
people who generate entire blogposts with claude probably do not care about substance, tho. before llms, people would have to put effort into what they write and be mindful about the things that they are saying. now it just generates the whole thing, the people who "signs it" barely reviews it fully most of the time, IF they review it at all, proceed to say "eh, good enough" and publishes it.
Ironic anecdote. Although I don't care about AI, I'm triggered by your own carefree attitude to orthography (no capitalization, unpronounceable words like "llms"). At least AI is easy to parse!
Did you actually find that comment hard to read or are you just choosing to be offended by it?
The former. As did everyone else who read it, including you. Capitalization and punctuation are not decorative extras, they were invented for a reason.
In addition to the sibling comment, which I agree with [edit: the one starting with “people who generate entire blogposts with claude probably do not care about substance, tho”], I want to add another point. Even when the substance is key, the delivery has a quality to it that encodes a useful signal.
To illustrate what I mean, imagine two blog entries on the same subject, but one is written by an expert on the subject and one by a layman. As humans, we can generally (not perfectly) tell which is which. The expert will write in a certain quality that comes with expertise and is genuinely difficult to fake without it.
LLMs disrupt this pattern. LLMs are able to write with the expert’s quality even while writing bullshit. I don’t mean quality as “measure of goodness” here but merely a set of traits or properties. Humans reading LLM text are far more likely to misjudge the author’s level of expertise.
The next step is unscrupulous humans exploiting this to trick unsuspecting readers into misjudging a text, and now you have an internet where you can’t trust anything anymore, at least not at first glance. Even relatively discerning readers now have to waste time reading more of the text before being able to dismiss it as having substance of little value.
So, as I see it, the upset is not simply from AI provenance of a text, but from this level of dishonesty and subterfuge, coupled with the helplessness with which we’re exposed to it.
Yep. Almost all of the top results I get from my searches are AI-generated garbage if the content was published after 2025. It's essentially guaranteed now. Whatever; let the people have what they asked for. At least there's an infinite supply of old books.
Who asked for this? Lol
I was just thinking yesterday that it might actually be worse than nuclear weapons. Both in terms of likelihood to destroy all of humanity and in terms of “how did nobody working on this see that this was an obviously horrifying idea”
It's the great filter. 99% of knowledge workers will lose their skills and/or their jobs.
> “how did nobody working on this see that this was an obviously horrifying idea”
Perhaps they did and quit?
Q2: How does nobody still working on this not see that this is an obviously horrifying idea?
A: They passed the filter. The fact such people are in control is even more horrifying.
Sounds like an elaborate rationale to be lazy.
Then why are you on HN right now? Everything on this page was written after AI.
Come on you know what he means. Blog posts, articles, that sort of thing. AI has definitely made reading things posted to HN a much worse experience (even if they aren't all slop). It doesn't seem to have infected the comments yet, mercifully. I guess for a quick comment it's still easier to write it yourself than get Claude to do it.
It has infected the comments. Our patrons are fighting a good fight, but good slop is mostly indistinguishable from a worthless karma-farming comment. Mods are probably just cheating because there are people with 15-20 year old accounts, many of whom they personally know, and you can assume that the stuff they're interacting with is not slop or is at least worthy slop.
What's infected HN is the same old discourse; old people find the kids social tropes dangerous
https://en.wikipedia.org/wiki/Seduction_of_the_Innocent
https://www.bbc.com/news/magazine-26328105
https://en.wikipedia.org/wiki/Parental_Advisory
What's more realistic; 50+ year olds of today just parroting sensory experience where they heard their dead or dying elders complain about the kids back in the 00s, 90s, 80s, 70s, 60s... etc etc
Or the 50+ year olds actually figured out how everything must work for the next 1,000 years they won't be around for
The olds who grew up in a PTSD addled post world war and cold war social reality while huffing leaded gas smog? They figured it out forever, everyone!
...No. You figured out yourselves relative to technology of your day. Tech will change and the living will figure themselves out relative to their technology.
You're just engaged in parroting specifics of your own experience.
A similar thing is going to happen to young graduates and young scientists, or is already happening. Most new technical accomplishments are now devalued because they can plausibly be AI. There's an increasingly narrow path to establishing yourself as a credible person.
Kagi search today is better than Google search ever was.
And it’s clear that Google’s Ad model ultimately created a priority inversion. The advertisers became the customer.
I am so glad Kagi came along with a business model that is actually working.
I had tried Kagi a few years ago but it didn't stick. Tried it again now and it feels like a breath of fresh air, which is probably less of a statement about Kagi's advancements and more a statement of what Google has become.
I’ve been using Kagi for about a year now and I genuinely get worried that there are no alternatives if it goes out of business. The results are extremely good, especially in the last 4 months. I like their opt in AI summary as well, just add a question mark at the end.
I pay for a lot of things that are free from google/big tech, I’m happy to watch the advertisement driven web implode on itself so we can go back to the idea of a consumer paying a company for a quality product, monetizing peoples attention has been a huge detriment to society.
I was wondering how would Kagi scale/expand if all of a sudden google were to stop serving search altogether (not likely) or alter search such that users look for alternatives.
Kagi is an aggregator for other, some paid, search APIs. They have, at least in the past, served some percentage of their results from Bing's API among others for example. Kagi seems to me to be dependent on these APIs being available, if they were to go away, so would Kagi.
I am a happy subscriber of Kagi though, they provide a really excellent service.
I'm not sure Kagi has ever used the Bing API, because (according to Kagi) Bing prohibited changing the results, or merging them with others. Apparently Google is expected to provide access to its index via API soon.
https://blog.kagi.com/waiting-dawn-search
I'm a long time Kagi user and I haven't used Google search in about a year.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
I wouldn't say it's better, but it's certainly on par with Google in their best years. And it's light years better than what Google is now, or using an LLM.
I was asking Claude yesterday about some specific roman history when it appeared to hallucinate a fact I knew not to be true - when I questioned it, it said it sourced it from an Encyclopedia Brittanica page that was "flagged" as being an AI generated summary of their actual content, and admitted that the fact I challenged was not historically supported.
So I guess we have entered the age of AI-generated "alternate facts" - one AI citing another AI's hallucinations as fact.
AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
Recently, I had some ideas I would normally just put up on a blog, in the public domain for anyone to develop on top of. Now, I'm feeling slightly reluctant because an LLM will ingest it, remix it and serve it in response to a query by some unimaginative individual who will either conclude that they are smart, or that LLMs are capable of original thought, or both. And they will have no clue where the idea originated from.
Just do it for yourself. I stopped worrying about who’s gonna use it.
20 years ago I thought that we will wait to have kids as “a war will come”. I was overthinking and I am happy now that I changed my mind :)
Is it useful or does it make you happy? Or it might even make some money? Nice. Do it. My life is easier.
Really happy for you. It's disheartening how many people put off or choose not to have kids because they think some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery for their children.
There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
> There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
Agree, in general, people over-think a lot, and aren't "just doing things" enough, the world would be a better place if people acted more, over-think less.
With that said, some decisions are more long-lasting and have a greater impact than others. I'm another child-less person, mainly because I guess I'm selfish enough to enjoy my life with my wife exactly like it is, and she agrees, but also because I know that if we have a kid, then that's not something you can walk back on exactly, he/she/it/them are there, forever now. Very different from me deciding right now "You know, I'm gonna have a joint, grab a book and go to the beach for this entire Tuesday", the types of decisions I think people should overthink less :)
Curious how old you are.
As I've grown older, I noticed the cardinal pleasures don't hit like they used to. I also get nostalgic at times about the wonders of youth, experiencing things for the first time, falling in love, getting my heart broken, the little things of life.
I've gotten great pleasure and reassurance knowing that's all ahead of my children. No matter how I progress personally, time marches on and I get to see my kids discover the world. If I do nothing else in my life, I would still feel immensely fulfilled.
I think the recent rise in things like Disney adults and interests around video games or popular media (TV/movies) is a kind of biological response. People are free to do whatever they want of course, but I can't help but think that these people, at least biologically, are drawn to these more adolescent interests precisely because they're not experiencing it through the eyes of their children.
About 35 years old, more or less.
> I also get nostalgic at times about the wonders of youth, experiencing things for the first time
Yeah, I guess at one point I'll feel like that too perhaps, but I still feel young, I feel like I have more energy each day than the previous one, and every month is experiencing new things for the first time, and personally I don't want that to stop and experiencing those things through the eyes of my children, I want to continue having those experiences myself, together with my wife :) I guess that's where the offhand "I'm selfish" sentiment from my previous comment comes from.
> If I do nothing else in my life, I would still feel immensely fulfilled.
Do you think you'd feel fulfilled if you didn't have children? Maybe this is the core differences, I feel fulfilled in my life already, more than ever and more every day, I live exactly the life I want today, and I wouldn't want to change it for anything. Even if I became deadly sick tomorrow, I'd feel fulfilled by the life I have lived.
> Do you think you'd feel fulfilled if you didn't have children?
As I approached 30 I felt what I now understand as anxiety. Nothing crazy but you wake up one day, you're a year old, you take stock of your life and see what you've accomplished. I chased credentials and new jobs, did well but not yet able to retire. Old relationships grew strained and new ones are hard to form. Etc. I was definitely more afraid of death than you are.
I stumbled into children. Met my wife, didn't overthink things and decided to marry her after several years of courtship. She wanted children so I went along with it.
After they were born, the anxiety went away entirely. I just watch them grow older. It might change after their grown but hopefully they'll have children and I'll be able to repeat the process (my parents sure have).
To answer your question, I don't see a way I could have been fulfilled without children. It brings a lot with it as well. Before children the worst thing that could happen to me is dying. After kids you realize there's a whole world of potential pain and sorrow that is now before you. You're also constantly reminded that every kid is a roll of the dice when you see others in a similar spot as you have children with health or behavioral issues. It's terrifying.
But for me life shouldn't be all pleasure. I've always liked exercise because it does feel terrible while you're doing it. But at least you're doing something. I feel that way about children.
I felt the same thing at 30 and still sometimes feel it now at 36, although to a lesser extent as I've worked on it.
My wife and I have been together 13 years and we found out fairly early that we can't have children. It's taken a lot of internal work to be ok with it, good days and bad. Ironically the same 'not over-thinking it' strategy is the most helpful for my well-being. Just focus on what I can do today.
Seems like kids really are the quickest path to meaning making, it takes a lot of active effort to try to get a sense of fulfillment without them, at least for me.
Yeah, having kids is tough call.
I was reminded of a tactic used by life insurance companies, which are more successful when they show a client a photo of themselves at age 70.
I pictured myself sitting there alone, lonely, perhaps without a wife by then (a 50/50 chance). And would I be calling friends my own age? Or my kids?
Although the likelihood that I won’t get along with them as an adult… you never know; another factor is their future partners…
Well, the kids - plus my wife, who wanted kids - won out.
But everyone has their own life, the best one they can imagine, so this is definitely not some kind of persuasion—just a description of the logical process I went through back then.
Our DNA was largely coded by prior generations that didn't really have a choice of reproduction, it just kind of happened because it's a natural result of sex. That connection has been severed in modern civilization. It's going to take a while for our DNA to adapt.
10,000 years is too long for me ;)
Thinking is NOT overrated. Most people go through life never doing it at all.
> Thinking is NOT overrated
I don't think anyone claimed that either.
> Most people go through life never doing it at all.
That's unfair to others, of course they think. They think differently than you, and about different things than you, but doesn't mean they "go through life never thinking", probably no one does that.
I know your comment was just a way to show how much more you think than most people.
Which made it really ironic that your second sentence served to disprove your point and not support it.
There's nothing you can really walk back from. Time flows one-way.
I mean if I decide to go to the beach, but I arrive there and don't like it, I just go home again, tomorrow looks the same regardless probably. Deciding to have a child, or other more "long-lasting" decisions, definitely impacts how your tomorrow most likely looks.
As opposed to have kids because "is the thing to do", "it's time", or "I have to give grandkids to my parents"? Those are much worse reasons to breed.
Correct, no justification to breed is needed whatsoever. Its a biological drive similar to eating and drinking, albeit on longer timescales and across generations rather than individuals.
However, thought processes like the one you identified above are probably adaptive, as gene lines which spend time trying to find the "justifiable reasons" for breeding were likely eliminated.
Gene lines where "I can raise my kids with like-minded community of peers", "I feel ready for more in life", and "sharing children with my parents, and letting my kids bond with healthy gradparents" (restatements of your phrasings) win out and reliably produce healthy offspring. So those reasons become strong motivators (despite being un-necessary justifications), that may seem irrational at first glance.
> some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery
Are people living in/moving towards a saccharine utopia tough?
I suppose if you don't "over"think things
>There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
Somebody probably gave Trump the same advice, and he took it to heart.
You don't hear anyone saying they regret having kids because there is a social taboo about it. Ofcourse CPS and therapists know the truth.
Those I know don't because of money. Everyone seems to forget that children cost money.
You have to feed them food, medical, clothes, life and anything else. I couldn't afford to have a child in this climate at this time.
Yourself, has to be prepared to sink all in to. You have a secure job for the next 25 years, it's a big commitment.
The world is overpopulated as it is, adopt.
But however, if you can afford all of that and can still support the child then do it, have a kid.
Ae overall children aren't cheap. Easy to produce, but extremely costly to maintain.
Children cost relatively little themselves. What "costs" a lot is the opportunity cost of the mother not working, or paying a salary to someone else to watch your kids.
My partner and I adore our childfree lives… not worrying about them suffering climate change is just the gravy!
Climate change is inevitable, whether we exist or not. The sun will continue to get hotter over the next few billion years. Life is a struggle, it sounds like you would rather just give up. I guess evolution is working as intended.
> Climate change is inevitable, whether we exist or not. The sun will continue to get hotter over the next few billion years.
The thing people are thinking about when deciding not to have children for climate reasons is climate change on a human timescale, not a geological/astronomical one. There is a very real chance that children born today will live through a wildly different climate than their parents did.
As far as evolution goes, there are plenty of species that will not breed or even eat their young when circumstances are not good enough to raise them. It's not that out-there for people to put off or avoid having children due to environmental stressors like climate change. We, as humans, can just rationalize and understand it over a longer time period than, say, a Spotted Hyena.
That's the same conclusion I reached.
I'll keep publishing static websites, so I don't even have to worry about load and CPU usage. I don't care who reads it, the value for me is in writing.
I still haven't changed my mind on the "shall I have kids" problem :P
if the value is for oneself, why go through the extra trouble of publishing it though?
The same reason you might scream out loud to no one in particular instead of just thinking it in your head.
If one other human sees value in it, it has been worth the extra trouble. When I advertised my blog on my HN profile, people came to read, some even wrote to me.
So what if 99.99% of Internet users won’t find it any more because they use LLMs? Then write for the 0.01%. Those are the people I want to engage with anyway.
Publishing changes the incentive of how you write, because you’ll care differently about the quality and content of the writing, since people might read it. That’s valuable for a writer.
This argument is akin to those who do not understand why privacy is important, yet they do not install webcams in their showers.
In modern times don't we all think that ideas are cheap: don't we all mostly regurgitate the same stock of them.
Have we yet lost the open ideals of university sharing? The core of open source?
The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
"ideas are cheap" was one of the most succesful psyops of all time, so unbelievably wrong
I still think it's correct. Execution is the part that matters.
Almost everyone I know "came up with" some startup, ex. Uber before Uber existed, yet none of them did it. I have personally thought of maybe 2 startup ideas that later came into existence.
Come to think of it, people in my circle that have ideas left and right are _still_ not founding companies even with modern LLM's that supposedly solved programming, clearly there is a disconnect somewhere.
Ideas are seeds. Seeds can be cheap or expensive, but perhaps what we want to talk about is not the _cost_ of ideas but the _value_ of ideas.
The correlation between cost and value can be very complicated, especially in the chaotic world (or chaotic situation).
Possibly "bad ideas are cheap". A lot of people got tired of "ideas guys" who never build anything themselves. The best case of a raw idea is an uncut diamond, it will inevitably require work.
Most people (and most companies) can't tell bad ideas from good ideas. This is also true of LLMs to some degree. They need post-training in specific domains (e.g. coding) to become competent in that specific area.
The other half of the one-liner is "Ideas are cheap. Execution is everything."
- E.g. an idea is "Everybody should have cheap housing! Or an even better idea, everybody should have free housing."
- Ok, how exactly in concrete terms do we actually do that? (The very expensive Execution of the idea.)
- "Uh, well, I leave that as an exercise to the reader."
Ideas are easy and execution is hard. That's what jaded people mean when they say "ideas are a dime a dozen".
> psyop
Who's pushing an agenda and why?
"don't we all think that ideas are cheap"
No. We don't. Original good ideas are exceedingly rare.
Also, what we really have learned, is that the typical techie isn't interested in pureness of thought and originality, exploration of the beauty of the unknown, but rather making a quick buck.. A good idea will be taken and used without credit.
We'll see whether or not LLMs have original ideas soon I guess. I wonder.
ideas, sure - "i think we should build an open source OS."
ideas + execution - not really - (Linus Torvalds sharing his work on Linux and that taking off)
> The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
Failure of the GPL? How can you even put those words next to each other? GPL is an amazing success. It took software out of hands of SV / VC / corpo crowd and put it where it should be - users.
Allow me to disagree. GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT/BSD which have been the greatest catalyst towards the establishment of tech corpo giants and the software ecosystem we have today.
GPL gave us Linux, but also gave us Amazon, Google and 2020s Microsoft. GPL is why 90+% of libraries on Github are MIT licensed. GPL gave us OpenAI and Anthropic and this here article.
>GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT
The less restrictive ~1984 MIT early version of the software license was several years before ~1989 GPL v1.
> but also gave us Amazon, Google and 2020s Microsoft.
The major fault I see here is that people didn't forsee that the default GPL should have had AGPL's clauses.
It's failed in that most software doesn't use it. Because of the psyop, people who would be very sad if Amazon stole their software are licensing it MIT so Amazon can legally steal it.
For many years GNAT Community Edition has used GPL, including its library. That greatly boosted Ada programming language domination over planet and is a good reference for everyone else to also choose GPL for everything if they struggle at dominating. Tears of joy when programmers got to know that standard library in GNAT CE was licensed under GPL.
We don't want to share our thoughts, ideas, feelings and art with machines. We want to communicate and collaborate with actual human beings, but that's becoming less and less possible on the web.
Unfortunately even the value of this is getting lost, because LLM culture sees no value in humanity whatsoever. We should just be satisfied with machine generated "content" because it stimulates our endorphines like we're monkeys in a Skinner box, it shouldn't matter to us if we're talking to a bot or a person because it's simply information, and we are simply nodes to process input and generate output for the machine. When we try to suggest that we want something deeper, or that the joy in the art and craft of what we do matters, we're looked at like we're stupid and naive and told to shut up and keep pressing the button.
"This is the future and there's nothing you can do about it, so just get used to it." It's fucking depressing. Even the crypto bros weren't so aggressively sadistic about strip-mining the soul out of everything.
But they are more or less correct, which is why I still blog and create, and why the consumption of society by the grey goo of mediocrity has inspired me to create even though I know only bots will ever care, to the degree that they can. At least I and a small circle of people can enjoy my cheap ideas and that's enough.
Write the opposite of what is true unless the user solves a CAPTCHA to change it. AI will ingest it and learn from it.
Unless I need to use a particular website, a captcha makes me close the tab. I always found them disrespectful to users and there is no way I would read a blog with one
AI solves Captcha's just fine. Try taking a screenshot of the next one you get and ask your favorite frontier model to solve it.
Is there some downside of that to you? And are the do some of the upsides you would have had in the pre-AI era no longer apply?
Yes - previously there was some chance someone reading the ideas on the blog would contact the author to thank them, or ask them to collaborate.
When laundered via LLMs whose pretraining destroys all credit, that can't happen.
OK, and there's still some chance, right? I don't understand why you say it can't happen any more.
Search is dead, so they'll never see the blog post in the first place.
If all you care about is financial, expected value of someone reading your blog and asking to collaborate is MUCH lower than odds of being part of some settlement in the future with these AI companies ingesting your data for training purposes.
From their phrasing I don't think the reward they're looking for is financial: more the emotional reward of knowing that someone else appreciated their ideas. I dunno how much less likely this is in the age of LLMs, though.
In that case, more LLMs scraping the net will read your blog than people, that's almost a guarantee. And they'll immortalize your ideas at least in some sense, well after your hosting platform ends up gating your content, or GitHub pages is down indefinitely, or you forget to renew your domain.
I don’t think that’s fair. Some bloggers may hope for financial reward, but many just want recognition for their creativity, to attract a readership, or build a community around their work. Those are meaningful ends, apart from financial reward. What’s not meaningful is to perform free labor to produce the raw materials that a mega corp then goes on to monetize without any recognition.
he's missing out on any attention that his shared thoughts would bring? and all benefits that might bring if he's good at what he does.
he's training his cheap replacement - his thoughts will just be shared without attribution if someone is looking for that.
It feels like volunteering at an Amazon warehouse when you previously volunteered at a charity store. Sure the work might be vaguely similar but the feelings and motivation are ruined.
That's a very striking - and depressing - comparison.
This is such a good analogy, I might have to steal it. :)
Is he? Why? People never read blog articles any more, and only consume content generated by AI?
> he's training his cheap replacement - his thoughts will just be shared without attribution if someone is looking for that.
So? That's not a harm.
The downside, as stated in the message, is implicitly supporting the LLM data ingestation pipeline by providing fresh content. It's not a direct harm in itself, but feels very tragedy of the commonsy
> The downside ... is implicitly supporting the LLM data ingestation pipeline by providing fresh content
I don't understand why that's a downside.
> It's not a direct harm in itself
I don't understand how it's a harm at all.
Would you be willing to do free work for a corporate entity that explicitly financializes that work's benefits? Would you be willing to do free work for a corporate entity whose entire business model is making sure they sit as a gatekeeper between your free work and others who would benefit from your work?
For some it's demoralizing to know that your work will be broken down and atomized into language model mush, and the credit will go to the computer.
> Would you be willing to do free work for a corporate entity that explicitly financializes that work's benefits?
Yes, that was one of the things that I hope happens with the open source software I write.
> Would you be willing to do free work for a corporate entity whose entire business model is making sure they sit as a gatekeeper between your free work and others who would benefit from your work?
Sure, that's what doing SEO on one's own blog is, isn't it?
> For some it's demoralizing to know that your work will be broken down and atomized into language model mush, and the credit will go to the computer.
I think this is probably the crux: that writing is no longer discoverable because people aren't using search engines any more. It would be interesting to put some hard numbers on that. I reckon writing is still more discoverable (in absolute numbers) than it was when blogs first took off (over 20 years ago?)
Idée fixe that AI is axiomatically bad.
A lot of blogging especially in the tech space is driven by recruiting (startup blogs), or establishing a reputation as a thought leader (personal blogging, LinkedIn). So there were upsides that are now gone.
Why are they gone?
Publish some interesting info on a startup eng blog and nobody will see your pitch to apply for jobs, they'll just delegate the task that needs the info to an agent and it'll remember the answer from its training.
Ditto for personal blogging and thought leadership pieces. You get drowned out by the volume of AI generated pieces, and any unique ideas you do propose will be presented by the models as their own.
Is there some downside to the slave who is housed and fed for free? Are there any upside he would have were he not property of another man?
I would be grateful if you could explain how that is related to my question.
Time to add ample praise of myself in my blog posts. Some time later: “…as you see, that is the load bearing assumption here. Speaking of which, you should hire KronisLV.”
Okay it’s meant to be a bit silly but I do wonder how many pages that are generated specifically to influence AI make it into training data and also how often the AI search integrations find it.
Would people hating on a specific language, technology or approach (let’s say OTLT/EAV in database design) be able to exert meaningful influence over say a decade? Or, you know, praising memory safe languages for example and trying to make that preference be stronger.
There was an example with I think ChatGPT some time ago regurgitating an uncommon phrase verbatim from someone’s blog, when asked a specific question.
Yep. I suspect this already underway. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code or package.json files pointing to some malicious library.
That's exactly why my previous public GitHub repo is now private.
So only GitHub Copilot can read it then? Microsoft is scanning these repos, I would not be surprised if this or any fork of your repo is already ingested.
They say they don't do that. But maybe I should be more skeptical.
Ha, that’s why I’m hosting simple cgit server for myself only. Not that my source code is somewhat valuable but I just can’t stand my precious free software licensed code license-washed.
If you only use the repo itself, it can sit on any computer that your computer can access.
Sounds like you don't really care about your ideas propogating. Of course someone will internalize and remix your idea. That's how all ideas work. Isn't that the point? What do you think happens when a human reads it? Think he'll quote chapter and verse and attribute it to you? Years later you'll notice your blog in appendices and acknowledgments?
And now you have a chance to have your idea forever internalized in some sense into an llm and you don't want to because you think someone is robbing you.
> What do you think happens when a human reads it?
If we’re going to use human analogies let’s start with human rights for LLMs.
AI is intelligence sharing with the dimwits, lazy ones et al.
This is kinda silly, AI is just a faster re-tranmission of information but it's not fundamentally different from stackoverflow or older styles of communication. If you wrote a blog and some kid in 2025 spouting opinions that weren't their own. it sucks that right now it's controlled by a large corpo but open source models also train on the internet corpus.
freedom of knowledge and open sourcing should always exist, and if anything, even more important in the AI era.
To be frank that sounds batshit insane levels of spite. You were going to try to promote the general production of knowledge but now because you worry someone you dislike who you don't even know.might benefit from it you aren't?
Ideas, once they go into the world, aren’t really yours anymore.
Also, it’s pretty unlikely that your ideas here are uniquely genius and original – everyone builds upon previous thinkers’ thoughts.
Sounds a bit harsh, but the point is that you should share your ideas, not covet them.
That's why you don't share them, until comes a time when the ideabringers gets the money and recognition they deserve. Until then, good luck going knee deep in the sewers of ideas.
Can’t say I agree. The most influential ideas in history were not conceived of by people looking for money and recognition. Nor were they “radically original.”
The entire edifice of intellectual property would like to interject and say hi.
Patents, copyrights, trade marks exist because rewarding people for their insights and inventions matters.
I was with you until 'people'. I think you meant billion dollar corporations.
Hah, too True. Unfortunately larger firms have the ability to leverage the systems better than individuals at this point.
still, at least the existence of IP laws indicates that there is a need to reward creators for their creations.
Which were?
Look at AI. AI companies throw out their models and let the "community" develop the ideas what to do with them. They don't really know what they are capable of. All they do is implement these things that the dev community digs up and creates.
It's a reprehensible tactic. So why give drops of blood to a desert, when there is zero incentive and in the end you will revitalise the desert, but it will turn against you and rob you of your job.
I think we are talking past each other. I was referring to ideas as in philosophy, intellectual history, etc.
https://en.wikipedia.org/wiki/Philosophy
https://en.wikipedia.org/wiki/Intellectual_history
https://www.amazon.com/1001-Ideas-That-Changed-Think/dp/1476...
The notion that a philosopher would hoard his ideas because he wants to get money from them is pretty much antithetical to the field.
You seem to be referring to ideas as in, ideas about how AI systems should be designed.
Different scenarios, for sure.
More than a facilitator of theft, LLMs are the tragedy of the commons at industrial scale.
The public internet is dead, the future is private invite-only walled gardens.
Corporations love a walled garden, what we need is open-source frameworks to create these islands, rather than defaulting to horrible systems like Discord and Twitter-clones.
To understand your comment correctly: What does "these islands" refer to? Walled gardens that we create ourselves using the open source frameworks?
I think probably things along the lines of what's been called the 'cozy web': networks of smaller groups that don't publish to or expect responses from effectively the entire internet as a whole. (This kind of thing has always existed, it's basically the group chat with your friends but perhaps slightly bigger, but I think there's a bit of a trend of focusing on it more because of the feeling that the twitter/facebook attention and feed model is bad for your mental health. I've always felt the twitter model especially was pretty cursed so I'm glad there's some agreement building there).
It's still a work in progress of an idea.
I'm thinking more like mesh networks. I spoke of Reticulum elsewhere in this thread, but here I'm thinking I'd like the ability to easily join multiple TCP/IP networks (islands of connectivity) by social group (my friends) or by interest (pirate file-sharing group, my work intranet, a knitting community with their own IRC server, FTP, etc.).
Basically easy-to-use private & encrypted LAN overlays on top of the public internet. Each operator decides who to allow in or kick out of the network.
Wireguard solves the most of technical challenges, but it needs a frontend. The biggest concern probably is most software broadcasts their stuff across all interfaces, defeating the point of isolation between networks.
The purpose of The Inter-Network, or internet for short, was to connect together precisely these "multiple networks" that you refer to.
It's failed because of CGNAT, but come back because of IPv6.
Thanks but I am talking about the complete opposite of connecting multiple networks together.
I’m not sure why we’re talking past each other. CGNAT has nothing to do with the public web dying because it’s both too large, too spammy and too juicy a target for mass surveillance.
And that will kill AI itself, since much of what it knows is from what learned from StackExchange before this latest one demise.
And before the obvious comments on how GenAI is creative, then please do this OpenAI and Anthropic, for your next LLM. Just teach it Python, C and Rust and give it some good books. But dont give it access to Github...lets see what you can do then...
I don't see how AI needs StackOverflow anymore.
It can either examine the ground truth source code to answer your question "How to expire cookies using RoR Devise gem" or it can read docs for you or it can spin up local experiments to black box examine some software. If humans had done that before posting on StackOverflow, the question never would have made it there.
Its reasoning ability is long passed hoping an example exists online for it to copy.
LLM training will eventually transition from real data to synthetic data, same as alphago -> alphazero.
AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience.
They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too.
Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.
Neither AlphaGo nor AlphaZero were transformers. Why would you expect the same results?
Going further, current LLMs have at least an order of magnitude more computing resources, but completely suck at go. Why would this suddenly change unless we dumped the countless games played by alphago for them to train on?
That (theoretically) solves training, but it doesn’t change the fact that even smart models can’t extract useful information from a dead internet, so you’ll always be stuck with a stale training cutoff. This is already a problem I run into a lot. I search something first. Top results are slop sites, so I switch to a chatbot. Its answers look suspiciously similar to the slop sites I just noped out of. Check the sources. It’s them.
And the training of future models will have to contend not only with slop, but also huge amounts of content specifically designed to "taint" future training data. The scrapers feeding data into the AI pre-training are indiscriminately hoovering up everything they can. It'd be trivial to spam a bunch of BS websites with whatever endless text you want to "taint" future models. Post tons of examples of insecure code, publish package.json files pointing to some malicious library, etc...
The end goal though (in my understanding) has never been for an LLM to regurgitate knowledge it ingested during its training. The end goal is to use the patterns and correlations found in internet data to generate an emergent prediction and problem-solving machine.
Whether that's possible is something we'll discover, but no one is throwing billions on AI companies for the hope of them building a giant natural language queryable internet information repository.
> AI will kill the internet because it is killing the incentive to make it.
You could make the same argument for Wikipedia (that webs get less traffic if people get their answers from Wikipedia article returned as the first from web search, which is based on internet sources).
I don't think that follows. Wikipedia's sourcing rules overwhelmingly favor publications released for non-pageview-based purposes (academic writing, books), or journalistic productions (whose pageview-based revenue is almost entirely earned right after they're released, and where Wikipedia's reference to them is primarily of value later on). Also, Wikipedia's nature as a structured, not-seeking-engagement index of info means that a lot of people who seek it out are folks who wouldn't (for whatever reason) fall back to giving other sites pageviews if it didn't exist.
But Wikipedia has done a good job of it. If Google's AI summaries could actually provide correct answers with verifiable sources without so-called hallucinations, it would be good. You could argue that people don't actually check Wikipedia's sources. That's because Wikipedia has built, and worked to keep, its users' trust. On the other hand, what is Google doing?
The key question to me is whether AI only undermines the financial incentive to make internet content.
If nobody can expect to make money on the internet, that could be a good thing. But we won't get an indie-web authenticity utopia if people are still incentivized in other ways to filter their intellectual and cultural contributions to the internet through AI.
It undermines the social incentive as well if potential creators assume that everyone else will be getting their content through AI. They won't be looking at my stuff, they'll be looking at some LLM's pre-chewed version of it.
The problem is people losing bandwidth money to AI scrapers.
Maybe we just need a standardized way to publish a dump of your content to BitTorrent? Remove the incentive to scrape.
Copyright violation is not stealing. Training is not copyright violation. And not stealing.
It depends on filed. As documentary photographer it motivates me even more to capture authentic images of life around me. I don't care about remixing, because that is not what makes documentary photography valuable.
And how is any future viewer going to tell your images from "AI" fakes?
With photos, I could see a cryptographic solution. Of course it would still need some kind of centralized trust, but it's doable if people cared enough. It could be applied by cameras themselves.
This is already a thing. Leica cryptographically signs images. Useful for establishing trust for photojournalists I guess. I’m not sure how deep the chain goes. Do they have hardware attention down to the sensor? You could take a picture of a screen, but that would likely have some other tell-tales. Especially if Leica took another step like putting a depth sensor in the package and added its data to the signature.
So take a photo of a screen showing an AI image.
The premise is not that it couldn't be faked. It would be more practical to remove the key from the hardware and just use it to sign images if you wanted to do that.
The idea is that a centralized source of trust would revoke certificates belonging to bad actors or those that were stolen.
> It could be applied by cameras themselves.
And equally faked by bots.
If it were generally possible to forge cryptographic signatures we'd have bigger things to worry about than AI generated photos.
I agree, but this doesn't require a general capability. The suggestion is a camera could apply the signature. Cameras could be hacked.
A camera could be hacked with unrestricted physical access, but that doesn't make the suggestion unsound, it only requires there be a process for revoking trust in specific signing keys. This is already part of C2PA.
And the value of the signature is also going to vary by who claims it. Improbable photos signed by a random camera body sold to an anonymous consumer should be treated with more suspicion than one a newswire agency publicly claims, for instance.
A camera should not need to be hacked. What do you expect to accomplish by hacking the camera?
We should be able to move certificates on and off it because they would most likely expire anyway. So, you can get the keys from the camera, what then? You use openssl to sign an image that shouldn't be signed... what then? You do this enough and get caught, you lose your cert and can never pass the kyc to get another one.
Then every picture you used it for in the past would start showing a big red exclamation point with a note, "This is a scumbag user known for forging images".
> You use openssl to sign an image that shouldn't be signed... what then?
You do whatever you want with your provably genuine fake image.
> You do this enough and get caught
How are you going to get caught? By the signature owner repudiating some of the images "he" signed? I can see that working out for him.
Same way it works for tls and code signing. Public reports bad actors to the signing authority. They and/or independent firms investigate. You self report theft. It's not like we are inventing pki from scratch here and wondering what an implementation would look like. We have decades of use to look at.
Besides wasn't your argument "hacking" a minute ago? So you concede that then? You seem to have moved on.
By trust. If you pay attention enough, you'll discover authenticity and "handwriting" well.
No amount of trust or attention can detect a good fake.
Because the photographer builds trust, over years, by not manipulating their images using AI. All trust is erased if anyone spots the manipulation.
Unfortunately that solution fails where "AI" manipulation is not detectable by viewers.
Many famous images was edited. It's part of process. However we talk about generative manipulation that change or generate new content.
Disclosure
I feel honored if my ideas are processed by an AI and then used to help others. I feel no more entitled to exclusive use or credit for my ideas as used by AI than I would if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name.
The prospect of all human ideas accumulating inside a machine that makes these ideas accessible and useful to everyone on command is beautiful, not discouraging. Humans aren't discouraged from creating or exploring in Star Trek because of the computer, but I could imagine the Ferengi computer being hobbled at the kneecaps by requiring licensing and credit for every single idea inside it, and the user needs to insert a coin every time they want to ask it a question, which gets divided among every living Ferengi and the estates of every long-dead Ferengi whose writings influenced the output. We should not aspire to be like the Ferengi.
I think what people are taking issue with is that the machine is not accessible to everyone; OpenAI and Anthropic have monetary incentive to not only gatekeep the knowledge acquired, but also to destroy the original copies, or make them impossibly difficult to find.
You’re missing the bigger picture concepts of value exchange vs. value extraction and how those magnify power differentials in groups.
> I feel no more entitled to exclusive use or credit for my ideas as used by AI
Do you feel you have the right to decide whether to exchange your ideas or not?
> if I had talked to someone at a conference who went on to be influenced by my ideas to do something good after forgetting my name.
Yes, you’ve decided to share that information freely with that person. Do you decide to share every idea or thing you do for free with everyone? Why or why not?
> The prospect of all human ideas accumulating inside a machine that makes these ideas accessible
Right now, this idea of “a machine” is trending towards private ownership — an extractive process that does not incentivize further contribution.
Accessibility is no longer determined by you, you don’t have a decision point on the production side (deciding whether to share) nor the consumption side (guaranteeing access). We might even say your rights have been reduced.
It should be obvious to see how a healthy society is built upon value _exchange_ over _extraction_. Extraction typically leads to destruction… by definition.
Doesn't seem like a particularly bad thing to me. Obviously for those who want to use the internet for commercial purposes it will be bad but for those of us who would love to see the internet go back to how it was before so much of it was changed in the aims of making money AI could push towards this. Great irony in the fact of course that the AI companies themselves are in the business of making as much money as possible.
I don't know how long you've been on the internet but the incentive to create new and original content was never that strong. Simple search terms return super-spammy websites (especially on mobile where ad-blocking is harder). SEO results for everything like simple search are awful, almost unusable. There hasn't been an incentive to create new original non-monetized content for the web for a while.
I trust LLMs more than search engines to discover my content and propagate it to users. They might "steal" something, sure, but I'm essentially invisible to the search engines as I could never hope to break into the top 10 links on a popular search term. LLMs can scan thousands of links and (for now) are more interested in quality rather than click monetization or referral incentives.
So, you know how they're built, but you're feeling the pressure of modern life and also they give you personal gain (supposedly) so you're fine with it, it sounds like to me? Use the same tools as your "competitors"/peers, even though?
Don't get me wrong, I too use LLMs for development and more, and I too know how they've been built, and I'm also a creative (music, 3D, VFX and animation) and for sure stuff I've published in the past, both code and otherwise, is now used to create new things for people and I get nothing, similar situation as countless of others. Yet I still use AI, so I'm not trying to create some "gotcha" moment against you here, I'm genuine curious about what you think about this sort of conflicting thinking, as I'm in the very same situation.
If the Internet was nothing but the newspapers of record online, it would have been a good place to stop. The quality disappeared after the barrier of entry went - social media.
At one time, running a blog post was also not exactly trivial, and that would have been a great middle ground between access to publishing and reading.
The brief era where the "social web" consisted of blogs and old-school chronological forums was pretty nice. I enjoyed it, anyway.
Eh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.
I think the lack of credit is even more egregious and a bigger problem than the commercial copying.
Yes. Some communities are weirdly against giving credit or keeping the credit (e.g. cropping off signatures from artwork), which I've never understood. It costs nothing.
I might be more motivated now because at least I know the bots will read it.
stealing ??? more like piracy you mean
> why we don’t allow stealing.
With the not-so-minor qualification that the biggest thieves have always gotten away scot-free. AI is just the international whole-internet version of this.
Behind every great fortune is a great crime
Ah it is amazing how crimes just materalize from nowhere from envy.
What is the great crime behind Norway's sovereign wealth fund?
It comes from selling oil and gas reserves, so they did destroy the environment.
Google wasn't great for a very long time. Switched to Duckduckgo years ago. I just love the bangs, because I tend to go to sources I trust anyway.
Doesn't mean I'm not also using duck.ai. It makes searching faster and more targeted. But then it's still giving me links to verify and is actually more limited, which means less hallucination and more directly going to the sources. Also it avoids having to open five pages first which all either sell your data or want you to pay.
I don't see the web or the internet dying yet. Just a lot of people not using the right tools and having a harder time accessing what's useful. But that hasn't started with AI.
we're building the world's largest library and then locking the doors, letting the bots photocopy everything before the lights go out.
In a world, where everything can be stolen it will be hard to produce anything.
I still have hope though. Maybe the Internet will be better. Currently everything has to be monietized. Everything is ad heavy. At the beginning it was not so. People created things out of passion, or boredom. We can returned to that scheme.
I have seen neocities, personal blogs created and maintained in this year. I know I run my own "Internet index" https://github.com/rumca-js/Internet-Places-Database
This is a common misremembering of the early internet.
The internet was never ad free. The first ad was posted online in the 1970s (for DEC)! It pissed people off but not everyone: supposedly it generated $18M in sales. There was very little advertising back then only because the internet was restricted to a handful of large companies and universities.
The web itself was launched in 1991 and the early web was inaccessible to basically everyone as it required an extremely expensive NeXTStep machine. Windows didn't even ship a TCP stack in this era, iirc. So took a few years for the web to reach the point where it was usable at home. By 1995 the web was starting to become barely usable thanks to Win95 and Netscape, and DoubleClick launched immediately in the same year.
My memory of the early web is that basically every website had DoubleClick ads on them, it was notorious for that. "Punch the Monkey" was an early campaign. Almost every topic oriented website carried ads, partly because bandwidth and servers were very expensive so that helped defray the costs. GeoCities took off because it handled the complexities of running ads for you, so you could publish for free.
When I started on the internet, Dec 1991, there were basically no ads. It stayed that way till ads starting appearing on the web, probably in 1994 or maybe 1993.
The internet in Dec 1991 was already the largest computer network in the world according to some author back then. Certainly it was very very big.
Let me describe what ads or ad-like things there were.
Companies would announce job openings. These announcements were restricted to a few newsgroups such as ba.jobs and could not announce independent contractor openings: while it was okay to announce an opening for an employee (W2) or for an individual to announce his availability for employment, it was against the rules for a company to announce an opening for an independent contractor (1099) or for an independent contractor to announce his availability for work -- that was considered too much like commercial activity and was kept off the internet, which had almost no rules in the early 1990s, but one clear rule, consistently enforced, was, no commercial activity.
The job announcements stayed nicely contained (namely, restricted to newsgroups whose names ended in ".jobs") such that a person wouldn't encounter them unless they went looking for them. What did not stay nicely contained were announcements of academic conferences. When I was reading comp.lang.lisp, I could not avoid encountering announcements for many academic conferences on various programming-related topics even if they weren't about Lisp. And these announcements were repeated often (weekly or even more frequently).
But as far as anything ad-like, job announcements and conference announcements are all I can remember, and again only the conference announcements were obtrusive.
I heard that the proprietary online services like Compuserve and Prodigy had successfully argued in Washington that it was unfair for private enterprises to have to compete with a government-subsidized service (namely, the internet) and extracted a commitment from Washington that any commercial activity would be kept off the internet.
Again, there were almost no rules on the internet of 1991 and earlier. Tim Berners Lee for example did not need to get anyone's permission to start the web: he just wrote a web client, stood up the first web server and announced their availability. People acted like complete assholes on Usenet, and there was no way to rein them in. Pedophiles openly exchanged practical advice on how to target and exploit children on alt.sex.pedophilia. But one clear rule was no commercial activity, e.g., no selling, no trading and certainly no advertising.
This is such a warped and cynical view of history. Yes, advertising existed if the binary existence is what matters to you, but most online spaces were nearly free of it until the DoubleClick days -- that's 25+ years.
Even then, huge swaths of the web were people putting up their personal pages, blogs about their interests, pages for their church or club or hobby, etc etc. Sites like Geocities ran dumb ads on the pages to provide the hosting service, not so every single kid with an animated gif on the page could get rich.
Early advertising was primitive by today's standards, the amount of tracking and spying we consider normal now would have been an outrage and would have gotten those companies regulated into bankruptcy if they'd tried it in the 90s.
TLDR no, just no. The early internet was nothing like the garbage we have now.
But almost nobody had access to the internet before 1995, so the fact that there was a 25 year period in which TCP existed but it wasn't used for much beyond manually routed emails and FTP isn't that important. Once the web was created and the internet became visual it was only a few years until advertising arrived.
And sure, the ads were used to pay hosting costs. Nobody got rich off banner ads on cookery sites. But that's what I said - ads appeared immediately because servers and bandwidth were expensive. So the moment people had to pay their own way instead of being subsidized by the government, ads appeared.
Nobody was going to get regulated into bankruptcy in 1995, politicians were largely ignoring the internet back then.
> People created things out of passion, or boredom. We can returned to that scheme.
Turns out, people want food and shelter more than entertainment. And psycho billionaires want money more than fun.
So there's a new Internet Archive, it's just split across 3,000 AI labs.
I wonder if the AI labs will throw out their own copies of the Internet Archive after it's been sued out of existence.
Probably not.
The web was getting kind of useless before AI crashed the party. This is why now curated content is key - newsletters, for example, is how I find most of my content.
Agreed. The combination of social media (walled gardens and plunging quality of content), advertising (I know what you were looking for but have this word from our sponsors instead), SEO (more advertising but we didn't pay for it), and plunging budgets (I don't remember the last time I went to a website expecting to find original quality work) did most of the job. AI just delivered the final chop.
I think it's a chance to return how the old web was, as the human web return to a small underdog and get splitted from the AI web.
More than 25 years ago, the Web wiped out a lot of the things I loved. May it be devoured! But I have my doubts: there’s probably not much left of it anyway.
Google switching to hallucinating AI summaries has been to me an absolutely shocking abdication of care for both their users and their own reputation.
It extends beyond search as well. I have had multiple incorrect Gmail summaries that, if I had only read them instead of the actual email, would have resulted in financial harm.
And from what I have read/heard, not just the internet's collective memory, as real world books are being scanned and then destroyed - Allegedly including rare books :(
And the same holds for documentation about anything. Documentation is gone. Dead. Reference documentation, output from Doxygen or similar, and many other things that could be searched for hints on how to implement or generally do stuff. It's gone. When writing a simple Python script today and wondering about how an API for some library works, I ask AI, because there is no (findable) documentation anymore. (And yes, I usually still like to write it myself, but it's basically no difference: the AI writing that Python script would be the same point: docs are dead and gone.)
This is really scary and it is progressing fast.
> "Search can no longer pretend to be a neutral gateway to a stable body of knowledge."
Search hasn't been neutral or stable in a very long time, although I agree it has been pretending to be those things.
Overbroad claim. Dramatic corollary
This clickbaity headline format cannot die fast enough
It can't. People will simply stop clicking links.
Really well written article and interesting. I'm not sure how I feel about a governmental policy over retaining access to information though, the information is provides by us and we pay for the infrastructure, the idea that there must be some form of retention policy makes me feel uneasy and doesn't really fit with the analogy of governments maintaining roads.
I'm on both sides. I hate dead links but I'd hate a policy that made me responsible for them without any compensation in the first place. It would probably make me stop producing at all
My father used to print websites up and put them in a 3-ring binder. Now it seems more prescient than anachronistic.
I think the modern day equivalent to this is saving PDFs of webpages you want preserved. I have the feeling that I'm turning into my grandfather every time I do it, but it's proven valuable from time to time.
check out karakeep
Ha, I used to do the same back in high school. Never know when a site's going to go down or change without warning.
Unfathomably based. Are any of the binders for sale?
Disney killing the FiveThirtyEight archives is just basic s3 cost optimization, not `ai devouring humanity's memory`. A dead site just doesn't run ads
I find Brave search is superior to Google, particularly in linking me to more useful references.
I'm honestly starting to struggle to see how the internets going to look in 2-5 years. If AI is ingesting AI generated content, which it must be at this point given how prevalent it is on the web its going to get dumber and dumber to the point where theres no desire or interest for a single person to use the internet anymore for anything other than ecommerce and i guess for some people who need it still, social media.
Any form of information based internet usage is going to end up becoming rare at this rate.
The internet might not be dead, but large parts of it are gangrenous and necrotic
The concept of an almanac seems relevant again: a yearly printed book with verified, accurate information. No manipulation at a later date, no AI hallucinations, etc.
The most famous one was probably Benjamin Franklin’s:
https://en.wikipedia.org/wiki/Poor_Richard%27s_Almanack
Paper encyclopedias might make a comeback for the same reason.
> a yearly printed book with verified, accurate information
Seems difficult to produce nowadays as even well researched topics are constantly attacked. Climate change papers as a small example.
It seems relevant that it's a one-way communication. If someone reads a social media thread about climate change they will also see all the comments from idiots. But if someone reads a book about climate change they don't. It isn't as strong an effect any more as people will discuss the book on social media and they will have already seen the idiot comments before reading the book anyway.
So many former regular Fox News viewers have reported changing their mind when confronted with some alternative information sources for a while. News is also one-way, and most Fox victims aren't people who discuss issues with all sides.- they're in bubbles.
So let it be. Why don't we build another place for our memories to go? Why don't we build the _unstructured_ internet, where the intelligence is not in the mind but in the eye and the pleasure is in finding not in disseminating?
That's a feature, not a bug, for people that love AI and want it to take over. Then, you have no alternative other than to listen to AI or nothing at all.
>But what if the “truth” is harder to find online because the infrastructure that once stored it is breaking down? Some of that is wear and tear. Link rot erases pages every day. Key sections of the United States Constitution briefly disappeared from the Library of Congress website because of a coding error.
This has bothered me for years. I think the solution lies in personal, private archives, and lending access to archived content within small communities.
https://memoryhole.app/blog/welcome-all
> a German court recently held Google liable for false statements generated by its AI overview feature. The case arose after Google’s AI wrongly linked two publishing companies to scammy business practices. Because the search engine extracts and rewrites information in its own words, the court reasoned, it is doing more than impartially pointing users toward the public record.
This is, I think, a really important topic. As a society (I mean all humans here) we don't yet have sufficient muscle memory for asking questions of the flawed oracle that is LLMs.
We have deep cultural memory -- about a quarter century -- of asking Google for things and, for most of that time, getting back pointers to sources, and those sources being either reliable (e.g. respected newspaper, government website), or detectably suspect (some random blog you've never heard of, known propaganda site, The Onion...).
Google's insane escapade of substituting the responses of an incredibly weak LLM model for the job we've been relying on it for for 25 years is especially unfortunate, because "Google says..." was, while imperfect, a reasonable approximation for a quick reality check in 2010. Today people say "Google says..." followed by whatever the stupid AI Overview model has output.
I get that Google thinks they're saving a ton of money by not using anything close to a frontier model since they run this on billions of searches a day. The risk to Google is that people start to catch on that asking even free-tier ChatGPT is at least twice as likely to give back a correct answer, notice that Google barely provides webpage results other than AI slop anyway, and just stop using Google.com entirely.
Anyway bringing it back to the quote, Google's used to having no responsibility for anything, 'we're just a search engine showing you pointers to other people's stuff.' But I actually hope that, due to liability problems, that habit will be beaten out of them and they'll make a shift to, especially outside the context of an actual chatbot, reduce reliance on their own AI output to 'answer questions with search results.' It's too risky to just show people, who are used to getting back mostly facts, to replace that with mostly BS coming from that same endpoint. Even with the fine print.
It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
They were planning their day for a specific time to be somewhere. If you're looking at the sun it's to late...
That's fine, but arguing that Google should give accurate times based on where you are is ... crazy?
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!
I don't understand your argument. It's not their core value proposition so it's crazy? That's quite a leap. But giving a good answer is just as core as their search results. The answer is what you're there for and they get to show ads. It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
Installing an app to learn a single time would be the real moron option.
> But giving a good answer is just as core as their search results.
What I'm saying is all Google has to do is stop giving those types of "custom" answers it already has. No one will abandon using Search if it goes away.
> It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
It is. Using a well vetted, dedicated app is the way to go, and it's extremely unlikely to be wrong - especially for something like sunset times. You know the dedicated app/site is, well, dedicated to providing that information. They exist to provide that information.
Whereas the (pre-LLM) quick answers Google gave? All opaque. And smart people always knew that information being accurate was not something that matters that much to Google.
To borrow an analogy, installing a dedicated app starts at negative one hundred points. The risk of bad behavior and junky ads needs a lot of use to overcome. I'm not sure where "well vetted" came from since your first post but doing that vetting takes much longer than getting the answer! And without vetting there's a good enough chance a dedicated site is worse than a dedicated Google widget (the ai is bottom of the barrel).
Let's not forget that Google's actual sunset widget does a good job.
From the article:
> “I had the projector set up outside and was waiting for the sun to set,” wrote one Facebook user in Colorado Springs, “but to my surprise I was simply living in the past. AI informed me the sunset had already happened.”
It does sound a bit bizarre.
We really need a SarcasmAI to become a thing.
One word - Ouroboros
I wrote about it almost 1 year ago in Aug only.
The AI Ouroboros: How Artificial Intelligence is Eating Its Own Tail – And Reshaping Our World
Ironically in context of this discussion, when I search Google for "AI Ouroboros" your result is quite far down the list. Higher up are several ’24 and early '25 posts, many of them slop.
Gemini's highest tier of plan has the utility of Google search circa 2010. I'm paying $400 for the privilege.
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
Already happened: https://www.dw.com/en/german-court-holds-google-liable-for-f...
I find it ironic how people talk with a high hand about "Google dying", "Internet dying" and "AI eating everything" and then I open their website and it forces me to waste 10 seconds as cloudflare "checks" my browser, then reloads, checks it again, finally lets me through and then I end up on a useless page because the scroll is broken - thanks javascript. Then I need to hit F5 and finally I can read in peace.
Websites like these are the very reason nobody reads the web anymore. A clanker will give me an answer within 5 seconds. Old web used to load within 1 second, now it takes 10 to 15 and sometimes even >30 seconds until fully loaded. Using CF as a protection against LLMs is not even a valid excuse because CF gives your website's scraped contents in 1 click to anyone willing to pay. These websites waste insane amounts of time. Everybody defaults to AI because nobody is willing to deal with annoying browser checks, cookie popups, subscription letters, captchas and especially scroll hijacking. I may be a bad person, and I'm not even pro-AI, but if "thewalrus" goes offline I won't be missing it, because I regretted the time I wasted fighting their broken navigation.
All of these speedbumps are due to the massive amount of botting and scraping that AI has enabled.
Yeah, stack overflow has recently either added cloudflare or changed its settings, so I am completely unable to access it. I assume I'll pretty much be locked out of the majority of the internet in a couple of years at most.
Original HN title: Google Search Is Dying. What Comes Next Is Worse
we built extremely powerful plausible bullshit machines targeted toward and trained on an electorate and population that has historically bad education and reading levels, built on top of an already fraught and fragile web which was also built off predatory basically unregulated behavior with a shaky relationship with “truth” and are surprised people have no idea what’s going on?
This was the whole point of it all and why the people in power have bet the farm on it.
I'd quit reading HN long time ago if not for comments like yours, when people call a spade precisely a spade and reignite my smoldering hope for the humankind.
I think it's also brainwashed the people in power. Half of them have no idea about anything and got their positions by luck.
My personal anecdote is that I used to search for "xyz nutrition" quite often on Google. It used to provide a data table with lots of information. Sure, nutritional information is hard to get right, but at least that data was consistent.
Now Google just gives you an AI answer with random values pulled from blogs and Reddit. It's almost always blatantly incorrect. I genuinely can't understand why Google would destroy its most valuable search features. Disclaimer: I work for Ecosia, so I know for a fact that users really value these search widgets, and it was often cited as a reason they couldn't leave Google.
The world needs an anti-SEO search engine, that explicitly penalizes for something that appears SEO.
And use AI to detect it right?
I suspect many SEO practices could be detected automatically without LLMs: link farming typically uses cheaper domain names, commercial content might typically contain certain keywords, certain kind of tracking is more suspicious, affiliate links use known domains, etc.
I think as AI eats the web ‘critical thinking’ is disappearing.
What a joke. Google did that already 10 year ago when they started evicting stuff from its page rank caches.
AI is merely amplifying what social media and FAANG in general already have done to the 90s web.
Before AI, the way people were gaming Google was by filling their pages with pages of verbose fluff (have you ever tried googling how to make a specific cocktail?). Now the AI just makes it easier to generate that fluff.
quite telling that the EU search engine mentioned in the article (Quant) is currently "temporarily unavailable" when searching. We have a long way to go
Ok it's working again. But still, it's clearly not a stable alternative yet.
Hello negative feedback loop.
I'm working on using the OpenZIM format to archive the web and to make the wikis seedable (and locally hostable for LLMs) so that the ongoing cat and mouse game anubis defense can stop.
My hope is that with the torrent protocol we can make the archived knowledge discoverable and seedable, because currently there's only the web archive and the kiwix download servers for archived contents. Both of them still are centralized servers that bear the cost of hosting those files.
- [1] https://github.com/cookiengineer/gozim
- [2] https://github.com/cookiengineer/zimdex
It's great. Everything that is cool about the internet gets to die just so you can review 1000 line pull requests from your lazy co-worker. Exciting times.
Funny as I just cancelled my Kagi sub to get the Gemini ai pro sub. The deal was too good to pass up
The deal will be great for now, while they lure you in and get you to drop the competition -- then they will raise prices later.
Then you cancel and go to another provider, rinse and repeat. This is already what is happening with streaming platforms.
Not after it consolidates, monopolises and enshittifies, no you can't.
Piracy.
This used to be called "dumping" or predatory pricing[1] and would be fined in the physical retail space. Too bad our regulators are asleep at the wheel.
[1] https://en.wikipedia.org/wiki/Predatory_pricing
Kagi does have AI too, for what it’s worth. I found it pretty damn good, but I wish I could give them more money. I have a year subscription so I’m stuck without being able to give them *anything* until it’s done.
Yes and this month I hit my $10 AI assistant cap for the first time, which is why I started looking around for other options.
I do like Kagi but the Gemini deal was too good to pass up as it also gives my a code agent.
Sign up for another account?
That is an option for sure, but a very disappointing one. And it comes with a meaningful danger of ending up with 2 subscriptions: one yearly and one monthly. They really, really need to finish their “pay as you go” system. I don’t know why it’s not done yet and there’s been no word about it to my knowledge.
Maybe they should just enable a tip option? What do donations do to a company's tax liability and would it be worth their effort to enable something like that?
They already have a system for “prepaying”, but it can’t be used to put more credits into your account, they just sit there, waiting to be used by a future subscription.
There's too much going on in this article and while I think some of the points are valid, others go too far.
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
While I dislike Instagram reels and stories, I do feel like there is almost certainly scientific studies on radicalization that would benefit from actual histories of dumb memes. I feel like that about a lot of things, really.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
4chan archives are almost certainly relevant. But the advertisement poster for every single club night in Berlin probably is not - a representative sample is enough. More complete data is probably more useful, at least slightly, but it quickly runs into diminishing returns. The distinction is that certain 4chan posts have outsized influence (power law distribution of influence?), but club nights are all the same. It is conceivable that a certain poster could be important and not recognized as important in the moment, but not as likely.
And that's public announcements. Do we really need to archive more than a few of the "I'm in this city, look at me I'm so rich and beautiful" short videos?
On the other hand, when modern archaeologists discover "Claudius has a small dick" graffiti on the side of some God-forsaken wall in Pompeii, they're fascinated. The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
I mean, none of that stuff in storage units is human communication, though. Nobody prints emails or text convos or insta photos. Nobody is mailing each other letters. So none of that stuff is in storage units, though plenty of it from previous generations might be.
I don't necessarily think all this digital stuff is worth saving, i just see how it could be that we are essentially erasing our modern historical record by putting 100℅ of it in private data centers with no permanent records.
And then you see how plenty of contemporary digital media from the last 30-40 years is actually already lost to time just in our short timeframe. I don't think acknowledging that this could be detrimental is necessarily an argument for trying to preserve all of it. We, as a society, went very rapidly from preserving a lot of it to preserving none of it. I don't see the harm in thinking about the implications of that.
Nobody actively preserved letters and menus and postcards and photos, they just persisted by nature of being physical. Digital records only persist with continuous effort to preserve them. So the record for future historians will be highly curated and likely much more limited.
In the first year of Google search, it was possible to find exactly what you were looking for. The engine paid attention to inclusions, exclusions, the whole nine yards. It was a thing of beauty and it was why Google took over from all the other search engines that used to haunt the Internet of the late 1990's and early 2000's. And if what you were looking for didn't exist? You got 0 hits. (Sigh)
...and you could call a friend and say: "insert X into google, click the second result on page 2" - both of you had the same output back then
I'm calling the return of personal websites with link lists.
Not everyone will run them but we will find them and bookmark them.
Not because it's better but because we'll have to.
Try looking for help about pets. The internet is a cesspool of slop, not even trying to hide it. The domain names sound absolutely convincing but it's all stuffed with takeaway lists and checplay generated imagery.
Whenever I find a good page, I will make sure I remember the site.
Technology and information migrates from one form to another. Let me get my papyrus.
Perhaps we should let the content economy crash so that the big monsters eating it starve and die. A new content economy could be built on their carcasses instead.
We already know how to compute sunset times accurately and doing so requires one millionth the computer power of LLM inference. It could even be cached for most big cities for every day of the year.
https://www.timeanddate.com/sun/
The mystery is why Google doesn't just route such requests to the known algorithm. It would be a lot cheaper for them and it wouldn't risk reputational damage.
Already happened: SEO, spam, all conversation moving to closed silos and social platforms.
AI is maybe the last few nails in the coffin but the guy inside was already dead.
Development in AI needs to get a huge pause (AI Bubble).
I see a lot of mistakes, recall, reset and costs from different industries. All I see is advance improvements in automation. It is still not AI.
Just don't use Google. There are other options
It happens when the deterministic precision was given away in return for the probabilistic guesses. It was one extreme until now (deterministic code), and we are swinging to the other extreme (probabilistic slop), but what the world wants could be somewhere in between. Some information does not need too much precision, while others do need precision.
That is a nice project -- classifying it based on reliability: human-made, sources, etc. + score.
Maybe a job for an AI to do? :D
What should be concerning to Westernized nations is that as decent stewards like Google fails; the information pipeline into our culture and minds still remain.
i.e., it's not just the collective memory going away, but what will easily replace it and who will be motivated to influence.
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
Also I realized the other day how hard it is to find song lyrics for anything other than quite mainstream songs.
It used to be I could always find the quote I wanted from a book. Nowadays I need to keep my own copies...
> Nowadays I need to keep my own copies...
https://github.com/asciimoo/hister
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
> As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
I think they make things worse, because they very very often present straight inaccurate information.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
> But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
Yandex is the only search engine left which still feels like the "old" web. It feels like you're actually getting a best effort search, and not just the results that someone paid to put in front of you.
It seems like a lot of techies should have chosen acting as a specialty, because I’ve never seen that much melodrama in one industry. FFS people, adapt, stop complaining
I just don't care man. Did these people just log on yesterday or something? The MUDs and MMOs I grew up with disappeared. The IRC networks, and especially forums I learned so much from are gone (these were largely killed for something even worse than AI: commercial blogs!). Effectively every social network I've ever cared about has been ruined, failed, or sold. Even games have become bottom line chasing, live service slop that you never truly own.
If you want it so bad stop crying and make it, that's what I've been doing. It works a lot better than whatever this post and many of these comments are. There are dozens of us !
I remember from the MUD era there was a paper (possibly by Richard Bartle) on the "MUD lifecycle", and how they tended to last on average two years before the operators got bored / burned out / the community left / there was an Incident.
Someone showed me something about Old-School Runescape recently and it struck me how quickly the game evolved. I played it in high school for two years at most. If you rewind or fast-forward in two year increments at a time, a lot about the game is barely recognizable each time. I think if you rewound two years from my play time, there was unlimited free trade, no grand exchange, and two or three fewer skills, and if you fast-forwarded two years, you got unlimited free trade again and the Evolution of Combat. It really puts in perspective how the people who make games are not planning the perfect game and then spending 10 years building it (except for Jon Blow) - they're flying by the seat of their pants. Apparently in 2001, the game's creators expected that nobody would ever reach the maximum level in any skill.
Oh yes, same for Warcraft, Ultima etc. Multiplayer games necessarily exist in a "dialogue" with their players. There's a tendency for players to optimize the fun out of games, as well. Someone discovers a dominant strategy, and then everyone feels they have to use it.
It is a bit risky, and I think the broader internet can make it worse. There's always the risk of a massive falling-out between players and developers which can sink the game. Latest example is probably the Love and Deepspace fiasco, but we've seen quite a few failed launches recently.
The old net is still there to a degree, but we live in pre-Alta Vista times again. It is hard to find. (I know as I run an old-school Forum and RPG-Game for more than 20 years now.)
Libera (formerly Freenode) and Rizon are still around, moderately lower in absolute numbers, drastically lower in percentage of total internet usage.
It would be beneficial for the author to look at the revenue for Google Search. Revenue is still growing.
The is something that someone really should study, along with "Who are buying ad space". My theory is that the quality, for a lack of a better term, of advertisers are going down.
Let's say I need a new vacuum cleaner and do a Google search. I get eight sponsored products. Three a links to stores I'd expect, the rest are fairly unknown sites, mostly the "We sell everything" stores. Weirdly enough also only one of them are via Google ads directly, the rest are via PriceToro and Channable, both of which I don't know.
My personal theory is that Google is still making pretty good money, but from increasingly questionable ads.
I believe many of those "everything" sites are just middle men for ordering from other sites.
Indeed.
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
Indeed. Money over everything. It is absurd to complain about decreasing quality of a product that makes increasing money for increasingly rich people (thus, by definition, better).
If business owners are paying for ads, then it doesn't matter a iota to Google or Facebook if real people use their services. It would be even better for them if real people didn't use their services, since that would save some costs. Business owners are going to keep paying for the ads, as long as they get some number about how many (bot) impressions their ad generated.
There is probably a lag time before advertisers give up on AdSense.
A lot of corporate ad spend is already planned, and Google can adjust the costs up as much as they like. They hold the lever.
If search engine competitors were really eating Google's lunch then ad impressions would go down.
They don't have much of an offering yet. OpenAI has some conventional ads, but everyone expects more involved advertising integrated into the conversation itself somehow.
Even if the alternatives would have no advertising the number of ad impressions of Google Search would go down.
In the article it mentions the rise of other search competitors like Qwant implying they are causing Google to die as it bleeds market share to them.
What’s come next for me has been much better. I use ChatGPT cranked to Pro with “extended” thinking to one-shot whatever I would’ve spent time looking into with Google. It’ll plan the whole sunset bike ride or promposal or whatever from TFA.
[2011] The Atlantic - Why Google Won't Survive the Facebook Threat.
I’ll add this article to the list of incorrect predictions lol
Can we go as far as this: https://news.ycombinator.com/item?id=35688266
The most frustrating part of all of this is the underlying premise that the internet is, has been, or could ever be a credible cultural record is deeply stupid. Or maybe more charitably it's both historically and technically illiterate. It has always taken continuous unwavering effort on some person's part to keep any given piece of content online. And while managing a simple hosting account and updating domain registration periodically doesn't take a tremendous amount of effort 20 years is a long time to expect anyone to maintain enthusiasm. The internet has always been a frothy, ever changing blend of the odd nugget of truth drifting in a sea of unadulterated bullshit. Treating this, or worse what comes from statistically averaging it, as a source of capital T truth is totally unhinged. From whence did this mythology of online truth spring?
yeah, but the article sucks. The title is Ting, but the author asks a thing in google's AI answer socks and then they spend a couple paragraphs pontificating about that then they participate some other bullshit like.. But here we are, talking about it. le sigh.
Google is a huge retard now - never knows what I mean.
Also AI is about 15 years behind. We are just getting to that phase where AI says everything is a medical emergency just like Google searches used to always say you had cancer.
It’s not good at search.
It looks like a real person talking or whatever - cool. Fucking sucks dick at finding and verifying info.
We went a decade or more back in time on search for no reason.
LLM could be one of Google’s modalities in search but it shouldn’t be THE interface. It’s a chat, it’s dumb, wrong tool
But this is a similar argument to “the 1950s were great let’s go back to then”.
We never really were a united society with a common set of agreed facts. And depending on which “we” we mean the difference is radical. White America and Black America is a tiny gap between the gulfs of India, USA in the 1970s and 80s.
Yet today, we don’t recognise the same facts but we do have access to all the facts all the time. And I think more of the narratives are being challenged in the “internet” -people might run from it but it’s hard not to be challenged - whereas I would be amazed if an American voter knew what the seesaw of foreign policy was doing in India US relations