Glancing at the Touhou 4 decomp[1], it's a lot better quality than a lot of AI decomps I've seen. It's matching, the variables are named sensibly, comments are sparse and comprehensible, and there's not much in the way of unaddressed Ghidra jank. And only in a month! My main complaint is that the file structuring seems more optimized for AI use than for mirroring the original intent of the devs. (Compare to this human-made Touhou 6 decomp.[2])
I know the retro game modding/decomp scene is already getting hit with a flood of low-effort ports.[3] Now they're going to be hit with a flood of half-decent decomps effortlessly generated by anyone with a $200/mo AI subscription, which become half-decent PC ports, which get modding support grafted on. It kinda just... totally eliminates that as a hobby entirely. I know people are going to be upset about this. And it's not even like with the AI art, where the pro and anti-AI camps live in their own camps. Solving the same puzzle twice just feels discouraging. Very similar to the problem academics are having with AI math/physics proofs.
I know a large reason people hate AI is because it's largely seen as tearing apart hobbyist/professional communities, eliminating reasons to collaborate with others and form relationships, and replacing it with an individualized dependence on a commercial product. It follows the same trend the tech industry took social media, from a method of connection to a tool of propaganda and paranoia and "@gork is this true?" This project, I think, is one of the most clear examples of that.
> Now they're going to be hit with a flood of half-decent decomps effortlessly generated by anyone with a $200/mo AI subscription, which become half-decent PC ports, which get modding support grafted on. It kinda just... totally eliminates that as a hobby entirely.
It really depends on your perspective of what the point of the hobby is. Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
Personally I think it's awesome that PC98 Touhou games are getting recomps. Being able to play the series natively on modern hardware, and all hardware for the foreseeable future, is great and preserves the legacy of the series. I understand that for anyone who enjoys the actual act of reverse engineering it feels like cheating, though.
I think people will get over it. Music is a good example. Drum VSTs and piano VSTs are going to sound way better than you unless you're a 99th percentile musician playing on a 99th percentile physical instrument. But people still play the piano and drums, because it's fun. Same idea; the existence of AI doesn't stop people who love the act of reverse engineering from doing it.
> It really depends on your perspective of what the point of the hobby is. Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
I think your question would make sense if these were the only two reasons. Someone who just wants to make sure their old games can run would see these tools as a good thing. People who just enjoy reverse engineering by hand wouldn't care this exists, or maybe even, a helpful tutor.
So perhaps then, the sentiment being as negative as it is, is revealing something about the hobby: that many participants, considered the metagame of being the first to do it, to be "smart enough" to have figured it out, the "street cred" if you were, is what they are lamenting, but they just either lack the words, or aren't being honest with themselves.
> So perhaps then, the sentiment being as negative as it is, is revealing something about the hobby: that many participants, considered the metagame of being the first to do it, to be "smart enough" to have figured it out, the "street cred"
This is a very ingenerous interpretation.
I am a hobby hiker. My mountain hikes are so easy that basically anybody could do them: no street cred involved, it’s not even worth to speak about them. But it is important to me that most goals are only reachable by walking and putting in the effort: if a company offered a 20$/mo service that brings subscribers to the endpoint, carried in a luxury litter by robotized carriers, I would feel that my hobby has been devoided of sense.
> if a company offered a 20$/mo service that brings subscribers to the endpoint, carried in a luxury litter by robotized carriers, I would feel that my hobby has been devoided of sense.
As someone else who hikes: why? I do not understand. Do you not enjoy the hobby of hiking for the sake of hiking?
The views will be just as beautiful, the air just as fresh, and the feeling of moving my legs just as energizing, regardless of how others may decide to get to the destination themselves. Why would the experience of others have any impact on my own experience of a trail?
Might be stretching the analogy too far, but look what happens to remote places that become easily accessible. They become commodified, crowded, and polluted.
There is something about a place requiring effort to visit that makes it special, and it loses that quality when it can be trivially accessed. The air isn’t quite as fresh when there’s a parking lot there.
Thanks for framing this in a way that I've struggled to phrase.
At every turn it seems like a lot of intellectual pursuits are just being voided by this technology.
I've heard the "one level up" argument in software with the supposition being that the work's still there, it's just becoming more abstract... but even in software, this doesn't really track.
The injection of this technology into not just intellectual but artistic pursuits (since there are now commercial AI image and music generators) should be a red alert.
What is there to live for once everything you once loved to do has been optimized for maximum revenue? Where every time you do something for fun, someone will come out of the woodwork when you show it off to either tell you AI could've done it quicker, or that they made an AI version with a flashier UI or something?
> What is there to live for once everything you once loved to do has been optimized for maximum revenue?
if it brings you joy, what does it matter how someone else does it?
If you run for fun, what does it matter that someone else does it faster, or in a paid exhibition optimized for max revenue?
Do it because it brings you joy, there is no external dependency.
> Where every time you do something for fun, someone will come out of the woodwork when you show it off
See, this is where things get tricky, either the activity brings you joy, and/or the sharing of it brings you joy, and/or the showing off of it brings you joy.
For the first two, it is irrelevant what someone else has to say or do about it; ignore the noise. For showing off, yes, some activities are no longer worthy, so change is needed to find those that are.
Hopefully for a hobby, or for anything that you love to do, showing off is not a major criterion for why you do it.
Showing off can be the point. If I did something tricky and creative and got an interesting (possibly even useful!) result, why shouldn't I be proud enough of that to have a "hey, look at this cool shit I did/made" moment?
That aside, I've been lucky enough in my career to work alongside, mostly, people who have a personal interest in using software as a tool, and knowing their toolchains very well, be that a language, a command line, a system architecture, you name it. If you found something cool to do with a tool, you'd share it to disseminate the knowledge. That culture is very quickly dying off, I feel, because LLM-driven work encourages siloing to a degree I don't think gets discussed enough.
Lastly, using LLMs to write code is its own type of exhausting. I thought I could keep up my hobby by mentally checking out of work and working on projects by hand in my free time, but babysitting the LLM, reviewing the massive corpora of code and documentation it puts out (including those put out by colleagues), and the insane volume of context switches this technology encourages in the workplace makes me want to stay the hell away from computers alltogether if I can help it.
We are not robots and telling people how they should feel and how their mood should not be influenced by this or that is akin to saying to depressed people to just stop being depressed. The divide-and-conquer method you appear to be using is worthless for human psychology.
The cases of a niche culture destroyed by dilution are so common that there are essays around it (e.g. Geeks, MOPs, and sociopaths in subculture evolution), so it seems to be a very normal and human thing to happen. Blaming the individual members for failing to feel differently is not going to fix that.
But you still can do it without LLMs/AI. And even better you can be proud that you did something "by hand" without AI help. Do you know that you can still cooking, singing, programming and so on? You do not need to buy ready meals, generate music by prompts or let model to figure out solution.
It is a lot of gatekeeping involved for sure. As a lifelong musician, who probably spent a few thousand hours to learn piano, guitar, drums and music production, I was shocked when people could produce reasonable sounding music with a mouse click and a budget of cents instead of endless dedication and a lot of money. I felt totally devalued and stripped of my „street cred“ to stay with your terms.
That's not (just) gatekeeping. People try to internalize a positive self image, from positive internal and external validation. Skills can also be a source of status or meaning (usefulness to others).
Taking that away compounds with the crisis of meaning we already have to make people feel more alienated than ever.
Gatekeeping was especially evident in MCU programming and embedded stuff.
Now you get decent answers for free, instead of "how don't you know that, read the 1200 page manual, stupid!"
A little nitpick on the sound quality of piano VSTs: I don't know about drum VSTs but I don't think your piano VST is actually a good example, because piano synthesis is actually pretty darn hard.
I haven't followed this closely the past few years, but it used to be pretty obvious you were listening to a VST compared to a real instrument in actual playing. When you compare recordings, sample based ones sound awesome as long as you just try a single note, but they were obviously missing something in the sympathetic resonances (sounding worse than my upright piano which is perhaps a 20th percentile piano), and the more physics based ones had not so great tones.
I've actually been hoping that someone would try to apply deep learning to piano VSTs, even emailed a few researchers, but no reply. If something new has come up I'm not aware of, I apologize.
Have you checked out Pianoteq? It is physical modeling and sounds incredible. It even runs on a RaspberryPi of you want.
When it comes to real instruments, Roland ships physical modeling in some of their instruments. They are a mainstream company.
So I think physical modeling is there but harder to get right than recording samples of piano keys being pressed. The results can be been convincing and it's nice that it requires little memory compared to sample based techniques.
> Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
It’s interesting to dwell on what ‘we’ appreciated before versus what we appreciate now. Since LLMs lowered the bar, it feels like raw human effort has come to be held in higher esteem.
Are we going to re-evaluate what we were impressed with in the past based on that new metric? For example, will accomplishments by experienced engineers be devalued in comparison to accomplishments by outsiders? Will the software engineering equivalent of ‘folk’ or ‘outsider’ art become more valued? There will be a natural instinct to distrust outsider contributions because, well, how did they get so good?
Are some metrics of human effort going to be de rigueur going forward?
I think you misunderstand the OP, the issue isn't the work if AI can do it we should still pool together and not waste tokens on the same problem 10-100x times because 100 people are doing the same thing in parallel.
Since it's not educational as it used to be, if you built your own version of n64 emulator number 10143 you got something out of it you learnt some interesting stuff, learnt software development, hardware etc.
Now we are literally all just taking a magic machine pouring money on one end getting half baked output the other but no progress is being made on money is being wasted. Since even if 1 person had built it everyone would have had access already.
This is the same wave of slop coding/vibemaxxing we had a few months ago where everyone was building product X for the 1 million-th time without any actual progress being made.
Math has a similar problem now, before if 1000 people tried to solve a problem they all grew from that experience, the LLM is not growing/learning anything from this effort.
All we are doing is burning resources for our pleasure, I say that as someone who has now written his own compiler, vcs and browser stack, with vibe coding, though largely I would claim I have used and tried to read the code, clean it up, to learn from it where different things break.
But even this feels absurd waste of resource imagine all of us collectively burning 200$ per month plans which assuming a at cost pricing we are all burning 200$ of electricity each month to do what exactly? Some folks are on their 10th subscription that means most people are burning 1000-2000$ of ultra cheap energy generally sourced from a mix of gas/oil turbines burning away at big data centers these days.
We could all solve it by making a site.. cough github.. and not duplicating work.. cough..
But somehow because it's easy no one seems to care.
I am hardly against software progress, and if AI gets better we will have better software (because most people write horrible software most of the time) so I am happy about it, to an extent that I can be, with my profession feeling somewhat challenged.
But given I have always worked as a fixer of broken stuff it's very much also feels fine to me. Either way I am very much against these 200-500 dollar plans that allow you to burn hundreds of dollars of energy for little productive gain duplicating the same work 100th time.
I am sure software developer using 100-200$ plans just for work can get more out of it. And it's probably making up for the effort they would have put in their work otherwise.
> It kinda just... totally eliminates that as a hobby entirely.
$800-level consumer 3D printers let people make highly detailed plastic models of things with very little level of technical knowledge, but people still buy and hand assemble LEGO kits for the fun of it.
When anyone can do these activities they will no longer confer distinction. So we are now motivated to find other ways to differentiate ourselves. This powers the next round of activities - our need to stand out by something.
It's a bit more than a need to stand out. It's a recognition that the definition of value has moved.
If a thing is no longer difficult to perform or to produce, then I might still employ or consume it if it's useful, but I will no longer value it. No matter how useful a potato is, it's not worth any money, time, interest, effort, emotional investment, etc. Even though in a vacuum it's almost like a magic ultra utility food, most useful and valuable thing ever.
Same for countless mass produced utilitarian goods like plastic bowls and basic computers.
I don't want be an injection molded cup or a potato not just because they aren't rock stars, but because they aren't valuable enough to be a person's identity or livlihood at all, even a basic unassuming workman's.
There is one valuable distinction to be made, ideas. There are only so many ways to grow potatos, IT in comparison offers unbound(?) possibilities.
The current projects we see are more showcases than original content, like remixes in the cassette era. Replication looses its value.
Why do you think it will destroy hobby? I think it will increase. When I was a child I wanted to have map editor or tools for modyfing closed sourced favorite games but for me it was not possible. Low level programming and reverse engineering is not for me (but maybe I will get into it soon) so I am happy about tools built for agents. Maybe one day I will decompile and extend favorite games without spending several years with or without success. LOL, I am not sure if I want to do it by hand.
And ihmo programming, math, doing music or anything else still can be hobby - with our without AI. Do not be worry, you do not have to use AI for your hobby stuff. You can still code, create from scratch crappy personal frameworks/libs and so on.
I only want for times when local LLMs/AI are powerful and available on average hardware, so I could use it for my personal stuff without monthly payments. I wish doing so many things on my own.
Fan efforts are breathing new life into older games, some of them locked away on older consoles or due to cloud tied services that the game developer has no interest in supporting any more.
Efforts like OldUnreal and the Peacock Project are just amazing. Recently, Halo fans have unlocked 128 player multiplayer modes using assets from Halo MCC and unlocked 16 player coop campaign mode. Microsoft has no interest in this cutting into their microtracsaction-strewn free to play game which isn’t as fun as older Halo.
Incidentally Ive started tracking really cool fan game projects here, including a disclaimer where possible if the project is AI assisted or not:
I've been decompiling DOS games recently with help from the $20 subscription. The one compiled by the Watcom compiler took like two weeks, with 2/141 functions still not matching. The one compiled by Microsoft C reached a match in two days more or less.
I think this points in two directions - first that $200 is really unnecessary for this domain.
Second, that with more experience we might be able to write a non-AI matching decompiler and then naming and simplifying could be done by people/AI. While I appreciate the hobby part of it, I think automated tools are the way to really make it reach everything with far less dependency on pockets. The benefit to preservation would be incredible.
> I know a large reason people hate AI is because it's largely seen as tearing apart hobbyist/professional communities, eliminating reasons to collaborate with others and form relationships, and replacing it with an individualized dependence on a commercial product.
For me personally that's a huge win, but it would be a shame if it turned out that the whole AI/anti-AI thing is really just single player vs multiplayer in disguise. Because then we won't reconcile it easily.
because someone ported aomething it doesnt mean u cant do it anymore... how does it eliminate hobbyists exactyl? hobbyists are the ones AI doesnt really impact because they pursue joy rather than financial gains in their craft...
Why would you dedicate months or even years to doing something just for the sake of it, when you know your creation will be buried underneath 20 other low quality versions of it made in a week by people who are doing it purely for attention and don't even care?
AI is just a tool, if your ability to form relationships was easily displaced by a tool, then I'd say it says a lot about the ability rather than the tool. It's the same argument of hand drill vs a drilling machine. No different. The ability to form relationships have always been with us. Instead of discussing in depth about code, perhaps it is going to be more along the lines of models, variants, self hosting, etc.
You can draw a line from Xenowhirl's disassembly of Sonic 1 and Simon Wai's research of the Sonic 2 prototype to the creation of Sonic Mania and the blossoming fangame / ROM hacking community, and from there the careers of a lot of indie devs. It's not that these people aren't capable of talking to each other, it's that they get to know each other in a personal and professional context by bonding over a project they have a shared passion over. If these disassemblies just popped up on GitHub one day, and if the fan games and the ROM hacks were created by solo devs with an AI model... well, we end up with quite a few neat projects. But we don't end up with the rich, vibrant communities that get created in the inherent process of collaborating on a project.
Of course, in a post-AI world people will still want to work together on things. But these sorts of tools do disincentivize it. That's the point I'm getting at here.
I think people who are doing these decomps aren't necessarily passionate about them to that extent. If they were a huge fan of the history of Sonic, they could use the resources now available to see how the game works, but they would also be internally driven to go beyond just a dump of the game's code on a repo. AI cannot decompile the social history of these games and the stories of the people who made them. It's just a matter if the person actually cares about that. Who knows what drives such an unusual amount of passion.
I feel a lot of these decomps were created by people not as passionate as your examples, because they are only showing up now, not in the pre-AI era. There is a flood of AI decomps only now because AI has lowered the barrier enough to where a project a person could not be passionate enough about to spend years of their limited time on Earth toiling on it is now easily accomplishable. It's a huge ask to put years of your life towards one thing. People might have obligations, work, family, children, more important hobbies, yet still wish one day they could get around to that one project, but couldn't find a way of prioritizing it over other things. I can't fault them for that. Now they can get it done in a month and still have time for other things that matter to them.
The claim is "we lose the rich communities that only exist from manual work." I still think the people who want those communities, like the people who still create art by hand in spite of the existence of diffusion, will keep doing their own thing. They were going to from the start, AI or not, but now that AI is a thing the separation is clearer to people on the outside, since they are seen to not rely on AI.
The communities that will appear because we can now organize AI decomp projects wouldn't have appeared at all with the barrier so high, had it not been easy enough to start them. The blocker was other obligations and not having the time to devote that was intrinsically required to finish them by hand, with the vast majority of people still being a member of outer society.
Dare I say it, maybe not all of the people into reverse engineering really enjoyed the process. It's just they put up with learning it because if you wanted to modernize an abandoned game in the pre-AI era, you had no choice but to learn those things. Only the people with the interest and patience remained, and those are the examples we are able to point to, while the people with only enough passion to get it done with AI assistance left a while ago and are only now reappearing.
> perhaps it is going to be more along the lines of models, variants, self hosting, etc.
Unfortunately those are also better discussed with LLMs.
I recently had to refresh on some algorithms and some manual coding for an interview where using LLM is forbidden. It was super fun, I forgot how fun solving software puzzles can be!
It’s unfortunate our industry has become so boring. Because our society values capital above all else, quitting my job to do something fun would relegate me to economical loser status. We really need a shift to the left in politics and a way to redistribute the productivity gains this technology allows.
> AI is just a tool, if your ability to form relationships was easily displaced by a tool, then I'd say it says a lot about the ability rather than the tool.
In abstract, sure, but that is a much shallower discussion that many people are strongly disinterested in. It's generic, where a human decompile may not be (as a cloner I haven't looked into those projects). I don't like the idea of a group of fans of art frames saying don't bother looking at what's in the art frame, our conversation is equal.
> It kinda just... totally eliminates that as a hobby entirely. I know people are going to be upset about this.
A rising tide lifts all ships, or however that saying goes. And a ship with a hole in its hull will sink to the bottom of the ocean. So don't sweat the surge of low-effort projects...
Your hobby is not going away. In fact, it's totally popping off right now. This is the perfect opportunity for you to distinguish yourself and lead by a positive example (whether that includes using AI tools responsibly, or not at all, is up to you). Just keep on hacking the good hack.
But the point here is that the rules and culture as they were no longer matter. They were there to ensure a level of quality that is now reached by an individual with a 200$ Claude subscription, so they should be discarded and the community will have to converge on some kind of new normal.
I'll have to take a look at this, but I've found that the top models do quite well at this work without any extra work.
The Windows Remote Desktop client has two bugs that have been driving me crazy for the better part of a decade. One day I got fed up and fed the binary to Claude, asking it to fix the two bugs. And it just did it. Patched one with some NOPs and adjusted a stack offset for the other. Produced a credible explanation for both, and in fact the fix worked. It helps that I was able to describe the bugs clearly, but still, it had to figure out the right spot among megabytes of executable code.
This is crazy, solving a bug from the executable is orders of magnitude harder than solving it from source code, this means that even with all the money they have, Microsoft didn’t bother to fix their shit.
It's harder in the exact way that AI is good at, though. You need to slog through very large amounts of boilerplate until you find what you're looking for.
It's not really more complex once you're practiced at it, just tedious.
Eh, there's a huge difference between having AI fix a bug and fixing a bug.
The latter implies that you've understood where and how it happens as well as noticed any other areas that may be impacted by fixing the bug or by not fixing it. AI just fixes it without thinking about anything else - which is great, but it explains why serious people (and Microsoft) aren't doing just that. We've spent decades worrying about quality, reliability and performance. Using AI is saying all of that was poppycock and it's faster to rebuild than to understand and fix. It's a different school of thought, and hopefully one that bursts down in flame in a few months.
Much harder than fixing from source, but much much much easier than making a bug report that Microsoft / Apple / etc will care about. You get better results from a wishing well than Apple Radar.
I am amazed by the jump in usability of models this year.
Some little thing annoying you about some software, some feature missing just point AI at it and burn some tokens.
(ok, maybe a few rounds of telling the model to do better are still needed from time to time)
vmconnect.exe (Hyper-V enhanced session) not going fullscreen after login when the window is fullscreen and you have to minimize and maximize to get fullscreen?
Have your agent debug the issue and come up with an elaborate hook system that leaves the Microsoft files untouched.
Codex/Chatgpt electron app coming with annoying or missing features (Mini, Invite friends, no mcp hotreload, windows updates failing, ...)?
Just point codex at codex and have it build an mcp Ouroboros to improve itself and enable dynamic js userscripts.
Deskflow clipboard sharing between devices missing file sharing, just have sessions on Mac, Linux and Windows collude to cook up a solution.
Yet writing proper documentation and reports that are human digestible still seems a bit out of reach, why would humans need to read anyway...
The novelty isn't binary patching itself. Its the the barrier to entry. Manually locating an bug in compiled code and calculating offsets used to require good skills and knowledge in the domain. Here, someone described the symptoms in a prompt (reminder, still plain English) and got an binary patch without touching a disassembler. To me, that's unbelievable.
I dunno, this isn't surprising at all now. The latest models are just really really good at reverse engineering. I asked Astra whether it was possible to do a task with a binary (commercial) library. I didn't even ask it to reverse engineer it but it happily went ahead and decompiled it, found some undocumented functions, figured out how they worked and then wrote an example usage for me.
My thought: Claude may not have fixed the bugs, but avoided the problem for the poster’s use cases. “Fixing a bug”, in my mind, means solving all problems for all use cases. And Claude didn’t have to do that.
The insertion of NOPs is a great way to remove a function call that might have crashed the program. If you didn’t need anything that function did, you’ve solved your problem but not fixed the bug.
By that definition scarcely any bug gets fixed ever. Perfect is the mortal enemy of good.
Also I don't think it's an honest take. You took this stance because the change was done by AI. If a person patched MS app he uses that had a bug that bothered him for a decade you would be more appreciative.
I wonder if this is why I've seen a lot videos popping up on my youtube feed related to vibe coded clones of various commercial apps in the past few days. Everything from clones of flagship products from Adob to Microsoft Office.
Using REA for something as high profile as what we're doing is likely to result in lawsuits. We're doing everything we can by the books.
We cannot look at Adobe sources. Use of Ghidra is disallowed.
REA is probably great for personal apps and for abandonware, but I think if you publish the results and it's found to have decompiled the original proprietary sources in discovery, you might be in for a bad time.
Using an LLM is not a clean room, imo. Its just IP laundering. Which is fine I guess if everyone is doing it, including the companies you're stealing from. I just dont know what the implications will be for progress.
Licensing/copyrighting encouraged people to think up of new things, and new ways of doing something. Now we're just all copying eachother.
What intellectual property is being laundered here if it's not even the same ecosystem (C++ vs Rust)? If it's so easy, and you just need to rebatch something existing, why has nobody done this over the last 20 years?
Meanwhile, Wikipedia has a policy that says screenshots of applications have to be "as small a version as possible" in order to meet the fair use exception for presentation of copyright work. I can only imagine an exact clone of the interface could be similarly ruled to infringe on Adobe's intellectual property.
Copyright is so flawed... This would mean an hommage is a breach of intellectual property in general, even if you paint it completely differently (say watercolor vs oil)
I agree. The vast amount of data ingested and internalized by LLMs has effectively been "laundered". But they are so powerful and evolving so fast that no one can be spared of their impact. We have to to learn to live with it.
Traditional proprietary software being "laundered" is just one part of the broader story...
Doesn't clean room typically apply to cases where the "dirty" team has legitimate access to copyrighted code, like the IBM PC BIOS which was published in the technical reference manual, and uses this access to write a functional specification for the "clean" team?
I'm not sure how clean room would apply to commercial applications distributed in binary form, as there's no way to look at even disassembled code without violating the license agreement and therefore being in breach of contract and subject to potential copyright infringement claims for copying or even continuing to use the software, let alone cloning it, and surely you're not going to be subject to a copyright claim based on familiarity with the application from merely using it.
The issue is that LLMs very likely have access to the source code of these products as part of their training data, and are then being used to generate clones. There's no barrier in the middle to ensure copyright violations don't leak
The reason why 'reverse engineering' has gotten so good is because what we're actually seeing is fully automated luxury plagiarism
> It is very unlikely that LLM had access to original source code of Photoshop.
It doesn’t matter. It’s likely had access to the knowledge of somebody who worked on that source code.
I have worked at companies that have done clean-room implementations. I was not permitted to look at their source or interact with those teams, because I had been exposed to the source of what they were re-implementing in a prior job.
As a human with a mushy brain I was never going to remember the source, but the fact I had been proximate to it was enough to lock me out. You have no idea what the LLM trained on, so you can’t prove a LLM derived clone is a “clean-room” implementation.
The first part (re: AI tools) is ostensibly not true. It all depends on the degree to which the data that they "don't train on" gets laundered and anonymized to the point where they are able to justify training on it without it being "yours" any more.
And I don't think that is publicly known at the current time. (If anyone has any tangible info on this, the please let me know...)
Well I hope they trust their LLMs completely, because I'm sure Adobe is currently going over their source code with a very fine comb to build a case. And they have LLMs to help with comparing too.
If AI models can generate designs faster and produce work that is "good enough", what is the actual future of design tools and the design profession in your opinion?
I've tested this myself with some frontier models and the results are kinda impressive enough that it raised the question if design skills are already obsolete. If that is the case, what use are these tools now?
I've seen a lot of "really cool unique designs" that are obviously just a digested regurgitation of amalgamated corporate slop. Anyone wanting an actual unique design language needs to hire actual designers.
Those designers may use AI tools but the tooling isn't to where they can be completely replaced. I challenge anyone to show me an e2e LLM design toolkit that can actually replace a designer, not just one shotted "wow that looks so cool" character designs.
Mightn't be so bad if all the value wasn't being captured by trillion dollar corps, despite them benefiting off the labor and creativity of countless non consenting individuals
Those trillion dollar corps also employ millions of people in very high paying jobs. And the extraordinary stock returns have made their engineers (and countless public investors) wealthy across the past half century.
People wishing for the death of the mega corps, are directly wishing for the obliteration of millions of great jobs.
>People wishing for the death of the mega corps, are directly wishing for the obliteration of millions of great jobs.
Yeah, I'll take the obliteration of a few million overpaid American jobs over the destruction of hundreds of millions worldwide, the social upheaval that goes with it, the destroyed families and the direct and indirect deaths. Fuck your promo packet and fuck your cozy desk job.
Probably not, there's relatively little secret sauce to something like Photoshop. It's just a lot of grungy work that, I guess, you can now delegate to an agent if you have enough money and time.
As an aside, I've heard a lot of hot takes about how this is the end of Adobe, but I'm pretty sure it misses the point. The main reason people pay Adobe is because it's a familiar line of stable, well-supported, interoperable, and actively-developed products. There's already plenty of cheaper or free alternatives (Davinci Resolve for video, Capture One / Darktable for raw, Affinity for photo editing and vector drawing, etc), and if Adobe survived that, I sincerely doubt they're going to lose pro customers to a vibecoded app where half the stuff is probably subtly broken or left as a TODO, and that will be abandoned in a week, because the whole point was to get that 1M YouTube views.
Not really. Someone using Adobe Suite of products cannot easily switch to Davinci Resolve or Dark Table.
These are NOT the same thing whereas this one particular suite of clone products is exactly menu by menu, keyboard shortcut by keyboard shortcut and feature by feature a pure clone of corresponding products to that extent that they even have a separate dashboard just to track feature parity: https://github.com/storytold/craft-repo-status-app
It is NOT there yet of course but it won't be there in a year? That's an absurd claim to make. I have checked the repos multiple times and every new release keeps fixing something that wasn't working before.
At some point Adobe will go after the clones, copyright their interfaces, close standards, patent stuff etc. The barrier for software used to be implementation effort, but now that it's out of the window, it will be more like other forms of IP (music, physical products, etc).
> that will be abandoned in a week, because the whole point was to get that 1M YouTube views.
Longbets 1 week, haha.
We're working our asses off on this.
Most of the team are artists who use these tools actively and we want the replacements for ourselves. I'm a filmmaker, so you can imagine my frustration of being bitten by the "unsubscribe fee".
I’m curious how much is your LLM bill. I’m assuming the idea is that LLMs will maintain the software? I’d be surprised, given the token economy, that you’d have to pay less than those subscriptions to LLM providers.
> I'm a filmmaker, so you can imagine my frustration of being bitten by the "unsubscribe fee"
What's the 'unsubscribe fee'? Are you referring to choosing something like an annual plan instead of 'Monthly' and only getting 50% back if you cancel partway through?
As you said you're a filmmaker, if someone chooses to obtain the rights to your film for a longer period instead of a shorter period and then changes their mind halfway through, do you similarly give them 50% back?
From what I can see, Adobe prominently and clearly displays from the first public checkout page the differences between 'monthly' and 'annual' options?
> From what I can see, Adobe prominently and clearly displays from the first public checkout page the differences between 'monthly' and 'annual' options?
It does now, because the FTC sued in 2024 and Adobe then corrected this, while settling in March 2026
You're doing God's work. Keep it up! PhotoCraft is really cool, and the pace of stabilizing progress is incredible! Don't listen to people who dismiss your work without getting their hands on it.
I am definitely interested in alternatives, but I am not interested in an alternative that isn't supported and maintained. Though I suppose at some point I can ask my own agents to support/maintain it?
No ethical qualms here but you guys should save your money and just use Gimp! It's already free, really capable, and comes from a good and pure place of making software for its own sake. That's the kind of thing that gets you users for decade, beyond any egui rust doo dads.
I've been really inspired by it. Ever since that sort of joking project came out, "Malus," that would clean room reimplement GPL software under more permissive licensing, I've been chewing on doing the reverse for proprietary software. I've been calling it the Manfred Macx theory of disruption, for the character in Accelerando that would patent and then release to public domain profitable ideas before a megacorp could lock them down.
I've got my wedding coming up so haven't had the time I want to dedicate seriously to finding some proprietary things to disrupt, but I've had fun using this as an opportunity to play with subagent orchestration, open weight models, various harnesses, and local models, to see what happens if I let some LLMs churn on it. That resulted in a hodge podge researched list of potential targets: https://github.com/508-dev/genairosity/issues
One thing that stands out is a lot of software kinda does already have a FOSS replacement, it's just not really how people want it to be. GIMP being the representative example. Photoshop people just don't like it, I'm sure for not entirely invalid grievances. However the maintainers of these kinds of tools are often strongly opposed to LLM involved contributions, again , often for not entirely invalid reasons on these incredibly complex projects.
So the long and short of it is that as powerful as LLMs are, we aren't quite where some of the doomsayers are saying we are, insomuch as proprietary software is dead. Even with reverse engineering we aren't there. There's genuine labor, time, and expertise moats around most of these programs.
Personally I'm shifting to trying to find niche abandoned software with no export flow. Even if I can only help out a couple hundred people, it sounds like a decent use of my time.
I built a from scratch rust RAW processing engine similar to the brains behind photoshop and lightroom - its much more powerful and its not even complete yet
LLMs are excellent at generating something that looks good, not so great at generating something that actually is good. All of those look really slick, but apparently the functionality is all broken
I see this liquid software being the unavoidable future. A computer that just does stuff in whatever way you guide it, in whatever way you like guiding it. The models will keep getting better and faster up to the point where everything is just happening in real time, no OS no drivers no programs, just an entity that can listen to you and can move around bits to accomplish whatever you are trying to do. Wanna post on HN? Any way that you can code such an action, the computer can just manifest it for you, on the fly, however you want it.
The expectation is of course that that will not be required. Competent models will become smaller and the compute to run them will become cheaper, once the market stabilizes. That will take some time of course.
Not today. But just like the average person can now afford a smartphone which is more powerful than a workstation not long ago, we will have enough computing power for good enough local inference, at an accessible price.
We're not in the 2010s anymore, consumer hardware is actively regressing. Smartphones are actually a great example to bring up, because I believe 2026 was the first year in which smartphone hardware did not advance at all (and arguably declined) relative to price, due to AI-related shortages.
Nobody can predict the future, of course, but I personally believe this pattern will hold. A decade from now, the idea of a consumer being able to purchase a personal device with more than 8GB of RAM will be a thing of the past, all significant compute will be done in the cloud using the massive amounts of hardware being hoarded by corporations as we speak.
I love such work. I hope to roll it into a package that anyone can use in the future. This work is only going to get more important because frontier models getting locked down will make this a lot harder over time.
For example, I use Claude as a bouncing wall for my thoughts and I pointed out that,
> GLM 5.2 was the only thing that helped HF while the agents were trying to access them. The "guardrails" stopped them from doing good. The Computer Fraud and Abuse Act exists. Courts exist. And computers and an internet connection have existed for a long time. There's also 17 USC 1201 provisions with the 1201 a 1 exemptions [Image #31] so in this case, a farmer should be able to work with you to access the tractor they own. Or... IDK... a kindle that's out of date? :) What is lawful and what isn't is rooted not within the act but within intent, purpose and mens rea. And this is something the law has been deciding for centuries now. At one end, your maker can't say that governments should decide while at the other end explicitly refusing to allow governments to be the ones who decide.
This was rejected for "Safety,"
> Opus 5.5's safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Edit and retry, or continue with Opus 4.8. Send feedback with /feedback or learn more: https://support.claude.com/en/articles/8106465
>
> Details: "[cyber]'
Note, the image here was the Library of Congress' page on DMCA exceptions.
Fundamentally, the idea that you can't reverse engineer things, make things, learn about biology or physics without permission is strange to me. These machines have been trained on the sum intellectual output of humanity, the global intellectual commons, and are being used to close off that commons?
I would be OK with their right to create such restrictions if they weren't lobbying the Government to restrict others, thereby ensuring that they control humanity's intellectual commons well into the future.
Perhaps I'm naive, but I think it's better for humans and the machines if we can all think, learn and build. But then again, I'm the kind of person who rejects the doomer pill.
I wanted to see if CVP approval changed this response, but it appears that with the release of Opus 5.5, Anthropic silently dropped me from the program, and has some strict new criteria in place to apply again, such as being credited for a CVE! I was only approved last month, too -- sad!
I have CVP with the new program (including mythos access) and still get constant denials for silly situations. Most recently I fed a URL to my agent from a security blog and asked if to add it to my obsidian vault with appropriate tags-- cyber flagged. You're not missing much. OpenAI and/or most Chinese models are much more lax in their restrictions.
This kills the talent pipeline, and it'll create a spam problem for the other folks because now people will try to github PR spam their way to getting on a CVE.
It's worth talking about the fact that you can't even talk about DMCA to a model trained on the Library of Congress unless you're one of the approved people. And that's before reverse engineering something or writing code.
So in this future, it sucks to be you if you're someone trying to make your small app more secure, someone trying to upskill, a tinkerer trying to bypass corporate lockdowns for a device they own (a recognized DMCA exception, btw), a teenager trying to learn about security...
It locks away much of the richness that produced hacker culture behind glass. You can look at their press announcements and PR pieces, but you can't touch.
And as they're lobbying the government for "sensible regulation," this inevitably leads to a future where computing is controlled.
It's the direction their existing reports are taking. They recently released one in September that talked about how they stopped "bioweapons." What were said bioweapons efforts? Oh, it was scientists using Claude for grant writing, paperwork and grammar. At national labs.
And this is being used to lobby against "dangerous" open-weight models because gasp a scientist might use them to write a grant! To make better antidepressants.
At what point do they start reporting someone taking apart an iPhone and trying to DIY a repair with a schematic as a thwarted "cyber security incident?"
A funny, but slightly chilling safety violation I once got was ChatGPT being unwilling to recite the full text of Article I Section 2 of the US Constitution, aborting as soon as it hit the passage about "three fifths of all other persons".
Another funny one was Claude's refusal to provide the original untranslated text of a passage from Dante's Inferno on copyright grounds, though in this case pointing out that no 14th century literature was subject to copyright anywhere in the world was sufficient to override its objection.
Agree on the talent pipeline. It can take a long time for someone to obtain a CVE that has their name on it. You don't start being a security researcher only once that happens.
Several years back, I was working on generating AVB2 hashes on top of modified Android distributions, to increase the security after an owner has made their desired changes. I was doing this before the age of LLMs. Among other things, this would've enabled the secure features to work again, and potentially reduce the risk of root access being usable by malware. But apparently I'm not a security researcher because I didn't get a CVE about it.
It's good that Stallman is still alive to experience this world. I think he never would've thought that freedom in software could come in this form, and as a long-time reverser myself, I've always held the opinion that his fixation on source code and the free software movement was not as liberating as it could've been.
"Source code? We don't need no stinkin' source code!"
Pretty sure Stallman would have some choice words to say about his idea of "software freedom" being twisted to include "relying on a corporation's subscription service for all your software needs and being unable to do anything yourself".
Please explain how this is more liberating than access to source code? Wouldn't this be reliant on continued and cheap access to the models? Even the open weight ones are heavily funded and not out of the kindness of their hearts. Just a few comments down is a thread about being blocked by the usual guardrails from the main providers.
But this is access to source code. All source code is now essentially open. If not now, it will be in a few months. You aren't any more dependent on the models then before , you just have the source code of everything that compiles on your computer. What happens in the future is anyone's guess, but for all software published until this point in history, closed source no longer exists.
No I mean sure, if you want to decompile it yourself from scratch, but once anyone does, then the source code is available to anyone. So from that point onwards, no one is dependent on LLMs to decompile it again.
really creeped out by the ai bro takes, specifically the ones "at least in a few months", almost as if there's no engineering principles anymore and just hype
Eh, decompiling is happening now, at current LLM capabilities. The resulting codebase compiles into something that is functionally identical to the original. What engineering principles exactly do you need beyond "works exactly the same as the original product", in this particular context?
I might be missing something but... How exactly is this better than telling Claude for example to "install and set up a full RE environment including Ghidra" on my local system and get to work? Like what does this do that my current RE methodology doesn't?
Probably it is not better. I think this is aimed at people who do not have a “current RE methodology”, do not know enough to specify things like Ghidra, etc., but who do have a desire to feel like they reverse-engineered and can reliably predict that a conversation with a chatbot will make them feel that way.
Everybody has their own set of skills and specific scripts and tools to do this stuff. You might use Ghidra as the the kernel of those workflows, but you still want something more than just Claude freestyling, at least for now.
I don't have reverse engineering experience, though I have all the prereqs to learn it. Anecdotally a few days ago I told my slow local Qwen3.8 in Pi harness to use Ghidra CLI to decompile a certain executable and it got entirely lost. Today I was linked to this, have it chugging along now, and it's making some sort of progress towards unpacking this thing. I don't know if it will nail it this run but it's a real improvement; I definitely have some useful info I could bring to another prompt or for myself if I cared to try manually.
But I just found an even easier way I should have thought of first - someone already dropped a reimplementation a couple weeks ago.
I don't know either. My only guess is that this has some helpful context for less powerful models, maybe.
We're kinda far into this LLM thing, maybe it's time to start selling harnesses by leading with how some examples were solved faster with this and how it saved tokens, or similar?
Because it's already been done for you? Sure, you can spend your own tokens on it, but it's probably cheaper and more time-effective to use something that already exists and does the job.
We're using vanilla Claude Code for ArtCraft apps [1], but we are especially careful not to touch Ghidra. We don't want decompilations or reverse engineering to spoil the work we're doing and expose us to copyright infringement.
Slightly off topic. I looked at REA's Android reversing support and found it still uses jadx mcp. That kills large scale APK reversing. jadx takes tens of minutes to preprocess an APK, build a code relationship database, etc. Even headless mcp is no exception. I basically can't use it to analyze APKs at scale, like 100 large commercial APKs in a pipeline, or 100 preinstalled APKs from a phone ROM to hunt bugs.
So I built droidasc. No memory bloat, no parsing slowdown. Analysis is in milliseconds. Global xref on a 300MB APK takes 1.5 seconds. I used it with codex to analyze phones from 3 different brands and found 2 RCEs and 5 root bugs in a few days. I plan to detail these at Black Hat Asia 2027. A friend used droidasc to scan various bug bounty targets at scale and found 10+ RCEs. Way, way faster than jadx.
APKPure is pretty good too, they don't have captchas for downloads so you can make a simple script to pull the APK from them based on the package name.
I’m surprised people don’t get refusals running this, or are they running it with cyber-enabled models like daybreak?
In my experience, even just hinting at reverse engineering to models from Anthropic or OpenAI leaves them extremely sensitive to refusals. After all, the same techniques used here can be used to find exploits.
I handed Astra Ghidra with some printer driver exes/dlls loaded, and a Windows 2000 VM with the software installed and said "reverse engineer this printer driver, it's hooked up and powered on on /dev/ttyUSB0, tell me when you're ready to print a test page". No refusals.
Really just a few attempts at cybersecurity related things (not even RE) with ChatGPT/Codex triggered a warning of breaking ToS by email. A few days later I was banned. No appeal.
My conclusion is that OpenAI will help you with cybersecurity and never really block or prevent you from doing the work, but then suddenly warn and/or ban you. Anthropic does the opposite, where they trigger safeguards all the time while working, but never really ban you.
Reverse engineering and decompilation is not currently blocked by the models, leading to the proliferation of LLM-assisted decomps/recomps. If you ask it to find bugs you're more likely to get a refusal, but simply pointing an LLM at a Ghidra MCP server and telling it to trace program flow, rename functions, or answer questions behaves like normal.
Ironically, I've had the opposite experience. My OpenAI account was permanently banned for doing too much RE work, meanwhile the worst I get on Claude is a downgrade to Opus 4.8.
I’m not the person you’re replying to, but for me just a few attempts at doing cybersecurity related things (not even RE) with ChatGPT triggered a warning by email. A few days later I was banned. No appeal.
The Chinese labs aren’t constraining their models as much as Western labs do. That’s why HF used GLM. If you remember the early times with how Gemini (Bard then) and Claude refused almost everything, you’d despise the idea of guardrailing. (Its important in some areas though)
This makes no sense to me. There are innumerable reasons to reverse engineer software that have nothing to do with security, some not only not prohibited, but in fact specifically authorized by law.
It'd make more sense to reject reverse engineering commercial software on the grounds that it likely violates the software's license agreement, and would therefore subject the user to potential breach of contract and copyright claims.
But this would also apply to uploading pretty much any non-self authored document to the LLM in the first place, albeit with fair use as a possible defense after the fact, so it still doesn't make much sense.
With C#, whenever codex need some info about the packages we use, it just use ILSpy to decompile the nuget instead of checking the docs. its much effective though. Even in my own packages, it use to gave back the bugs or gaps that i need to patch up.
Does anyone have insight into the sudden trend of game decomm / recomp and mashups like Minecraft in GTA, Mario in Skyrim, Cod in Life is Strange, Among us in Portal 2 and so on? What triggered it? Is there a specific new tool, technique or LLM that sparked this interest? I am curious.
Opus 5.5 being cheap and extremely capable + added YouTube virality.
I’ve been doing decompilation with LLMs for 4 months, and 5.5 is so insanely good I now can finish Windows 95-98 games in just a week. And by finish I mean get byte-to-byte exact output from your decompiled code to retail with clear names, semantics and no “Ghidra smell”. With older Opus models you would have to manually guide them and inspect ASM yourself otherwise they will just burn through your tokens (which also used to cost more!)
I have extremely mixed feelings about how this impacts things like game decompilation projects, but when it comes to cracking DRM my feelings are only positive. (Although I'm sure those on the other side of the DRM are unhappy)
When you crack some DRM, the DMCA makes it challenging to share your work with the world, regardless of the morality of your use case (say, repairing one's tractor). There is no longer a pressing need to share that work, when anyone can just say "computer, sync my spotify collection to my jellyfin instance, using correctly tagged FLACs", and it goes away and does it from first principles.
I don't know how much this matters as we get into a world where any software system can be composed in a few hours.
As the value of a particular piece of software decreases, the value of reversing its specific implementation does as well.
Reverse engineering is about discovering specific methods or protocols. Not for porting entire implementations to new platforms or products. A lot of the specific methods and protocols have been sucked up into the LLM weights, so the need to go digging for these patterns is dramatically reduced.
The biggest application I see here is with digital archaeological work (the opposite of new things).
Oct 5 ~4.2k
Oct 7 ~10.3k
Oct 8 ~17–18k
Oct 10 54.6k live
Going from ~4,226 to 54,614 stars in ~5 days is about +50k.
Even stranger is forks:
438 → 10,376 forks.
That is an unusually high fork acceleration. Current forks are about 19% of stars.
I’ve been trying to rebuild and modernise an old abandoned ms dos game, simply by pointing Codex (6.1 Sol) at the game directory, and it works insanely well. It’s able to understand the data formats, unpack graphics and sound assets, and reconstruct game logic.
> Install REA and connect it to this coding agent using npx rea-agents@latest setup. Show me the setup plan for approval, then verify the installation.
We have achieved the next evolution of installation by `curl | bash`!
A big difference between the safety of "next > next > next > finish" and "curl | bash" is one of them is dynamically loaded from an external source that could change between runs, and the other can be fully downloaded and vetted in a single check, and then once it's safe, it's probably safe 10 years from now.
Which is why we have sha1, md5 and sha256 hashes on display, so you can validate with a very high level of certainty, especially for sha256 at least for now, that the file is the same one.
We have existing paradigms for this.
Additionally, installers are signed with certificates on Windows.
All of these are strictly more trustworthy than curl | bashing.
Yeah installation has always been such a security issue. So many programs are just random links that download a file. You have to trust that the host has not been compromised all packages that were used to build it were not compromised etc.
With ai models getting better we may be able to do analysis on the actual underlying bytes of the files we download to properly scan them for malicious code patterns and build systems which sandbox programs and watch inbound and outbound traffic/ system level actions from them and flag suspicious requests for further analysis by smarter models.
REA shows that ai are very good at understanding low level code and reverse engineering it so this could potentially be applied to application level security aswell.
At some point I think it would be better to just redo the thing from zero like was done with Beyond All Reason.
The mechanics of BFME are the important part, not the intellectual property of the films. The way the units move, the resource system and power points are what make it special. Gandalf and Lurtz are simply wallpaper on top of something that is already very competent.
Now there are no excuses to reverse engineer the most notorious closed source binaries out there including from Nintendo's system software to CUDA, and nvcc from Nvidia and make it all "open source".
The only problem is the lawyers at all those companies will be readying their lawsuits, and given they have tons of money; they do not care and will come after anyone.
Why wait months for Nvidia to fix their compiler bugs when its far more faster to do it in the open? Libraries that invoke nvcc don't want to wait either.
So really closed-source compilers are a hindrance and you're subject to unspecified time-frames to even get to fix your issue rather than doing it yourself.
Intel and AMD have already open up theirs without question. Nvidia is the one who claims to be supporting openness in AI but can't even open up nvcc let alone CUDA.
I really doubt there's anything interesting in there, judging from experience of looking at implementations of mobile banking apps and websites. Horrendous API, some intrusive scanning/probing/fingerprinting code and usually pretty weird ass ad-hoc cryptography protocols + some mechanism to try to lock API use to OS vendor giving a blessing for its use, so usually Google has to say, "ok, you can call API of your bank to access your money" or whatever on every single use of the app. (which is one of their ways to increase moat around their walled OS gardens) Otherwise you're out of luck using a mobile app. Websites don't have this issue, yet. Mozilla doesn't gate my access to bank APIs.
Backend will have some shuffling of data around some ledgers + a lot of ceremony around auditing + shit ton of CRUD mess and arcane connections to other systems/institutions. Probably the nightmare of nightmares codebase, if frontends are any indication. :D
Same with healthcare ime. I'm more on the human services side, but I was reversing their apis since pre llm days when I was a relative novice. Many don't even minify, so you can step thru the frontend src in devtools.
And yes, if frontends are any indication, just seeing tip of the shit-berg
There are some days I just hit a raw debugger in AuthForRealThisTime() under a three-paragraph jsdoc written by bot which is itself under the commented-out Auth() function. So many questions arise, and I can burn hours rabbit-holing the stack
Glancing at the Touhou 4 decomp[1], it's a lot better quality than a lot of AI decomps I've seen. It's matching, the variables are named sensibly, comments are sparse and comprehensible, and there's not much in the way of unaddressed Ghidra jank. And only in a month! My main complaint is that the file structuring seems more optimized for AI use than for mirroring the original intent of the devs. (Compare to this human-made Touhou 6 decomp.[2])
I know the retro game modding/decomp scene is already getting hit with a flood of low-effort ports.[3] Now they're going to be hit with a flood of half-decent decomps effortlessly generated by anyone with a $200/mo AI subscription, which become half-decent PC ports, which get modding support grafted on. It kinda just... totally eliminates that as a hobby entirely. I know people are going to be upset about this. And it's not even like with the AI art, where the pro and anti-AI camps live in their own camps. Solving the same puzzle twice just feels discouraging. Very similar to the problem academics are having with AI math/physics proofs.
I know a large reason people hate AI is because it's largely seen as tearing apart hobbyist/professional communities, eliminating reasons to collaborate with others and form relationships, and replacing it with an individualized dependence on a commercial product. It follows the same trend the tech industry took social media, from a method of connection to a tool of propaganda and paranoia and "@gork is this true?" This project, I think, is one of the most clear examples of that.
[1] https://github.com/N0zoM1z0/th04 [2] https://github.com/GensokyoClub/th06 [3] https://www.pcgamer.com/gaming-industry/pc-ports-of-old-cons...
> Now they're going to be hit with a flood of half-decent decomps effortlessly generated by anyone with a $200/mo AI subscription, which become half-decent PC ports, which get modding support grafted on. It kinda just... totally eliminates that as a hobby entirely.
It really depends on your perspective of what the point of the hobby is. Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
Personally I think it's awesome that PC98 Touhou games are getting recomps. Being able to play the series natively on modern hardware, and all hardware for the foreseeable future, is great and preserves the legacy of the series. I understand that for anyone who enjoys the actual act of reverse engineering it feels like cheating, though.
I think people will get over it. Music is a good example. Drum VSTs and piano VSTs are going to sound way better than you unless you're a 99th percentile musician playing on a 99th percentile physical instrument. But people still play the piano and drums, because it's fun. Same idea; the existence of AI doesn't stop people who love the act of reverse engineering from doing it.
> It really depends on your perspective of what the point of the hobby is. Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
I think your question would make sense if these were the only two reasons. Someone who just wants to make sure their old games can run would see these tools as a good thing. People who just enjoy reverse engineering by hand wouldn't care this exists, or maybe even, a helpful tutor.
So perhaps then, the sentiment being as negative as it is, is revealing something about the hobby: that many participants, considered the metagame of being the first to do it, to be "smart enough" to have figured it out, the "street cred" if you were, is what they are lamenting, but they just either lack the words, or aren't being honest with themselves.
> So perhaps then, the sentiment being as negative as it is, is revealing something about the hobby: that many participants, considered the metagame of being the first to do it, to be "smart enough" to have figured it out, the "street cred"
This is a very ingenerous interpretation.
I am a hobby hiker. My mountain hikes are so easy that basically anybody could do them: no street cred involved, it’s not even worth to speak about them. But it is important to me that most goals are only reachable by walking and putting in the effort: if a company offered a 20$/mo service that brings subscribers to the endpoint, carried in a luxury litter by robotized carriers, I would feel that my hobby has been devoided of sense.
> if a company offered a 20$/mo service that brings subscribers to the endpoint, carried in a luxury litter by robotized carriers, I would feel that my hobby has been devoided of sense.
As someone else who hikes: why? I do not understand. Do you not enjoy the hobby of hiking for the sake of hiking?
The views will be just as beautiful, the air just as fresh, and the feeling of moving my legs just as energizing, regardless of how others may decide to get to the destination themselves. Why would the experience of others have any impact on my own experience of a trail?
Might be stretching the analogy too far, but look what happens to remote places that become easily accessible. They become commodified, crowded, and polluted.
There is something about a place requiring effort to visit that makes it special, and it loses that quality when it can be trivially accessed. The air isn’t quite as fresh when there’s a parking lot there.
Thanks for framing this in a way that I've struggled to phrase.
At every turn it seems like a lot of intellectual pursuits are just being voided by this technology.
I've heard the "one level up" argument in software with the supposition being that the work's still there, it's just becoming more abstract... but even in software, this doesn't really track.
The injection of this technology into not just intellectual but artistic pursuits (since there are now commercial AI image and music generators) should be a red alert.
What is there to live for once everything you once loved to do has been optimized for maximum revenue? Where every time you do something for fun, someone will come out of the woodwork when you show it off to either tell you AI could've done it quicker, or that they made an AI version with a flashier UI or something?
> What is there to live for once everything you once loved to do has been optimized for maximum revenue?
if it brings you joy, what does it matter how someone else does it?
If you run for fun, what does it matter that someone else does it faster, or in a paid exhibition optimized for max revenue?
Do it because it brings you joy, there is no external dependency.
> Where every time you do something for fun, someone will come out of the woodwork when you show it off
See, this is where things get tricky, either the activity brings you joy, and/or the sharing of it brings you joy, and/or the showing off of it brings you joy.
For the first two, it is irrelevant what someone else has to say or do about it; ignore the noise. For showing off, yes, some activities are no longer worthy, so change is needed to find those that are.
Hopefully for a hobby, or for anything that you love to do, showing off is not a major criterion for why you do it.
Showing off can be the point. If I did something tricky and creative and got an interesting (possibly even useful!) result, why shouldn't I be proud enough of that to have a "hey, look at this cool shit I did/made" moment?
That aside, I've been lucky enough in my career to work alongside, mostly, people who have a personal interest in using software as a tool, and knowing their toolchains very well, be that a language, a command line, a system architecture, you name it. If you found something cool to do with a tool, you'd share it to disseminate the knowledge. That culture is very quickly dying off, I feel, because LLM-driven work encourages siloing to a degree I don't think gets discussed enough.
Lastly, using LLMs to write code is its own type of exhausting. I thought I could keep up my hobby by mentally checking out of work and working on projects by hand in my free time, but babysitting the LLM, reviewing the massive corpora of code and documentation it puts out (including those put out by colleagues), and the insane volume of context switches this technology encourages in the workplace makes me want to stay the hell away from computers alltogether if I can help it.
We are not robots and telling people how they should feel and how their mood should not be influenced by this or that is akin to saying to depressed people to just stop being depressed. The divide-and-conquer method you appear to be using is worthless for human psychology.
The cases of a niche culture destroyed by dilution are so common that there are essays around it (e.g. Geeks, MOPs, and sociopaths in subculture evolution), so it seems to be a very normal and human thing to happen. Blaming the individual members for failing to feel differently is not going to fix that.
But you still can do it without LLMs/AI. And even better you can be proud that you did something "by hand" without AI help. Do you know that you can still cooking, singing, programming and so on? You do not need to buy ready meals, generate music by prompts or let model to figure out solution.
It is a lot of gatekeeping involved for sure. As a lifelong musician, who probably spent a few thousand hours to learn piano, guitar, drums and music production, I was shocked when people could produce reasonable sounding music with a mouse click and a budget of cents instead of endless dedication and a lot of money. I felt totally devalued and stripped of my „street cred“ to stay with your terms.
That's not (just) gatekeeping. People try to internalize a positive self image, from positive internal and external validation. Skills can also be a source of status or meaning (usefulness to others).
Taking that away compounds with the crisis of meaning we already have to make people feel more alienated than ever.
Oh, yes.
I must think of all the people who poured countless years of their lives learning all the C++ bullshit..
Gatekeeping was especially evident in MCU programming and embedded stuff. Now you get decent answers for free, instead of "how don't you know that, read the 1200 page manual, stupid!"
A little nitpick on the sound quality of piano VSTs: I don't know about drum VSTs but I don't think your piano VST is actually a good example, because piano synthesis is actually pretty darn hard.
I haven't followed this closely the past few years, but it used to be pretty obvious you were listening to a VST compared to a real instrument in actual playing. When you compare recordings, sample based ones sound awesome as long as you just try a single note, but they were obviously missing something in the sympathetic resonances (sounding worse than my upright piano which is perhaps a 20th percentile piano), and the more physics based ones had not so great tones.
I've actually been hoping that someone would try to apply deep learning to piano VSTs, even emailed a few researchers, but no reply. If something new has come up I'm not aware of, I apologize.
Have you checked out Pianoteq? It is physical modeling and sounds incredible. It even runs on a RaspberryPi of you want.
When it comes to real instruments, Roland ships physical modeling in some of their instruments. They are a mainstream company.
So I think physical modeling is there but harder to get right than recording samples of piano keys being pressed. The results can be been convincing and it's nice that it requires little memory compared to sample based techniques.
> Were decomps/recomps a hobby because people wanted to make sure their old games could run on modern hardware and be extended with mods etc., or were they a hobby because people enjoyed the actual process of using Ghidra and reverse engineering by hand, and the decompiled game was just a side effect?
It’s interesting to dwell on what ‘we’ appreciated before versus what we appreciate now. Since LLMs lowered the bar, it feels like raw human effort has come to be held in higher esteem.
Are we going to re-evaluate what we were impressed with in the past based on that new metric? For example, will accomplishments by experienced engineers be devalued in comparison to accomplishments by outsiders? Will the software engineering equivalent of ‘folk’ or ‘outsider’ art become more valued? There will be a natural instinct to distrust outsider contributions because, well, how did they get so good?
Are some metrics of human effort going to be de rigueur going forward?
Study former examples of 'AI' doing this : the Luddites and various artisans displaced by 'AI' looms and 'AI' factories in the 19th century.
I think you misunderstand the OP, the issue isn't the work if AI can do it we should still pool together and not waste tokens on the same problem 10-100x times because 100 people are doing the same thing in parallel.
Since it's not educational as it used to be, if you built your own version of n64 emulator number 10143 you got something out of it you learnt some interesting stuff, learnt software development, hardware etc.
Now we are literally all just taking a magic machine pouring money on one end getting half baked output the other but no progress is being made on money is being wasted. Since even if 1 person had built it everyone would have had access already.
This is the same wave of slop coding/vibemaxxing we had a few months ago where everyone was building product X for the 1 million-th time without any actual progress being made.
Math has a similar problem now, before if 1000 people tried to solve a problem they all grew from that experience, the LLM is not growing/learning anything from this effort.
All we are doing is burning resources for our pleasure, I say that as someone who has now written his own compiler, vcs and browser stack, with vibe coding, though largely I would claim I have used and tried to read the code, clean it up, to learn from it where different things break.
But even this feels absurd waste of resource imagine all of us collectively burning 200$ per month plans which assuming a at cost pricing we are all burning 200$ of electricity each month to do what exactly? Some folks are on their 10th subscription that means most people are burning 1000-2000$ of ultra cheap energy generally sourced from a mix of gas/oil turbines burning away at big data centers these days.
We could all solve it by making a site.. cough github.. and not duplicating work.. cough..
But somehow because it's easy no one seems to care.
I am hardly against software progress, and if AI gets better we will have better software (because most people write horrible software most of the time) so I am happy about it, to an extent that I can be, with my profession feeling somewhat challenged.
But given I have always worked as a fixer of broken stuff it's very much also feels fine to me. Either way I am very much against these 200-500 dollar plans that allow you to burn hundreds of dollars of energy for little productive gain duplicating the same work 100th time.
I am sure software developer using 100-200$ plans just for work can get more out of it. And it's probably making up for the effort they would have put in their work otherwise.
What's a VST?
Virtual Studio Technology. Often referred to as a virtual instrument or a (virtual) synthesizer, and works as a plugin in digital audio workstations.
> It kinda just... totally eliminates that as a hobby entirely.
$800-level consumer 3D printers let people make highly detailed plastic models of things with very little level of technical knowledge, but people still buy and hand assemble LEGO kits for the fun of it.
...and today even adults are served explicitly with all these LEGO editionsn of whatever building or classic car,for above500 USD+ :-D
When anyone can do these activities they will no longer confer distinction. So we are now motivated to find other ways to differentiate ourselves. This powers the next round of activities - our need to stand out by something.
It's a bit more than a need to stand out. It's a recognition that the definition of value has moved.
If a thing is no longer difficult to perform or to produce, then I might still employ or consume it if it's useful, but I will no longer value it. No matter how useful a potato is, it's not worth any money, time, interest, effort, emotional investment, etc. Even though in a vacuum it's almost like a magic ultra utility food, most useful and valuable thing ever.
Same for countless mass produced utilitarian goods like plastic bowls and basic computers.
I don't want be an injection molded cup or a potato not just because they aren't rock stars, but because they aren't valuable enough to be a person's identity or livlihood at all, even a basic unassuming workman's.
There is one valuable distinction to be made, ideas. There are only so many ways to grow potatos, IT in comparison offers unbound(?) possibilities. The current projects we see are more showcases than original content, like remixes in the cassette era. Replication looses its value.
Why do you think it will destroy hobby? I think it will increase. When I was a child I wanted to have map editor or tools for modyfing closed sourced favorite games but for me it was not possible. Low level programming and reverse engineering is not for me (but maybe I will get into it soon) so I am happy about tools built for agents. Maybe one day I will decompile and extend favorite games without spending several years with or without success. LOL, I am not sure if I want to do it by hand.
And ihmo programming, math, doing music or anything else still can be hobby - with our without AI. Do not be worry, you do not have to use AI for your hobby stuff. You can still code, create from scratch crappy personal frameworks/libs and so on.
I only want for times when local LLMs/AI are powerful and available on average hardware, so I could use it for my personal stuff without monthly payments. I wish doing so many things on my own.
Yep, totally agree with this sentiment.
Fan efforts are breathing new life into older games, some of them locked away on older consoles or due to cloud tied services that the game developer has no interest in supporting any more.
Efforts like OldUnreal and the Peacock Project are just amazing. Recently, Halo fans have unlocked 128 player multiplayer modes using assets from Halo MCC and unlocked 16 player coop campaign mode. Microsoft has no interest in this cutting into their microtracsaction-strewn free to play game which isn’t as fun as older Halo.
Incidentally Ive started tracking really cool fan game projects here, including a disclaimer where possible if the project is AI assisted or not:
https://fangamesdb.com
I've been decompiling DOS games recently with help from the $20 subscription. The one compiled by the Watcom compiler took like two weeks, with 2/141 functions still not matching. The one compiled by Microsoft C reached a match in two days more or less.
I think this points in two directions - first that $200 is really unnecessary for this domain.
Second, that with more experience we might be able to write a non-AI matching decompiler and then naming and simplifying could be done by people/AI. While I appreciate the hobby part of it, I think automated tools are the way to really make it reach everything with far less dependency on pockets. The benefit to preservation would be incredible.
> I know a large reason people hate AI is because it's largely seen as tearing apart hobbyist/professional communities, eliminating reasons to collaborate with others and form relationships, and replacing it with an individualized dependence on a commercial product.
For me personally that's a huge win, but it would be a shame if it turned out that the whole AI/anti-AI thing is really just single player vs multiplayer in disguise. Because then we won't reconcile it easily.
because someone ported aomething it doesnt mean u cant do it anymore... how does it eliminate hobbyists exactyl? hobbyists are the ones AI doesnt really impact because they pursue joy rather than financial gains in their craft...
Why would you dedicate months or even years to doing something just for the sake of it, when you know your creation will be buried underneath 20 other low quality versions of it made in a week by people who are doing it purely for attention and don't even care?
It’s more likely that the higher quality version will rise to the top, and the slop versions will be buried in obscurity.
AI is just a tool, if your ability to form relationships was easily displaced by a tool, then I'd say it says a lot about the ability rather than the tool. It's the same argument of hand drill vs a drilling machine. No different. The ability to form relationships have always been with us. Instead of discussing in depth about code, perhaps it is going to be more along the lines of models, variants, self hosting, etc.
You can draw a line from Xenowhirl's disassembly of Sonic 1 and Simon Wai's research of the Sonic 2 prototype to the creation of Sonic Mania and the blossoming fangame / ROM hacking community, and from there the careers of a lot of indie devs. It's not that these people aren't capable of talking to each other, it's that they get to know each other in a personal and professional context by bonding over a project they have a shared passion over. If these disassemblies just popped up on GitHub one day, and if the fan games and the ROM hacks were created by solo devs with an AI model... well, we end up with quite a few neat projects. But we don't end up with the rich, vibrant communities that get created in the inherent process of collaborating on a project.
Of course, in a post-AI world people will still want to work together on things. But these sorts of tools do disincentivize it. That's the point I'm getting at here.
I think people who are doing these decomps aren't necessarily passionate about them to that extent. If they were a huge fan of the history of Sonic, they could use the resources now available to see how the game works, but they would also be internally driven to go beyond just a dump of the game's code on a repo. AI cannot decompile the social history of these games and the stories of the people who made them. It's just a matter if the person actually cares about that. Who knows what drives such an unusual amount of passion.
I feel a lot of these decomps were created by people not as passionate as your examples, because they are only showing up now, not in the pre-AI era. There is a flood of AI decomps only now because AI has lowered the barrier enough to where a project a person could not be passionate enough about to spend years of their limited time on Earth toiling on it is now easily accomplishable. It's a huge ask to put years of your life towards one thing. People might have obligations, work, family, children, more important hobbies, yet still wish one day they could get around to that one project, but couldn't find a way of prioritizing it over other things. I can't fault them for that. Now they can get it done in a month and still have time for other things that matter to them.
The claim is "we lose the rich communities that only exist from manual work." I still think the people who want those communities, like the people who still create art by hand in spite of the existence of diffusion, will keep doing their own thing. They were going to from the start, AI or not, but now that AI is a thing the separation is clearer to people on the outside, since they are seen to not rely on AI.
The communities that will appear because we can now organize AI decomp projects wouldn't have appeared at all with the barrier so high, had it not been easy enough to start them. The blocker was other obligations and not having the time to devote that was intrinsically required to finish them by hand, with the vast majority of people still being a member of outer society.
Dare I say it, maybe not all of the people into reverse engineering really enjoyed the process. It's just they put up with learning it because if you wanted to modernize an abandoned game in the pre-AI era, you had no choice but to learn those things. Only the people with the interest and patience remained, and those are the examples we are able to point to, while the people with only enough passion to get it done with AI assistance left a while ago and are only now reappearing.
> perhaps it is going to be more along the lines of models, variants, self hosting, etc.
Unfortunately those are also better discussed with LLMs.
I recently had to refresh on some algorithms and some manual coding for an interview where using LLM is forbidden. It was super fun, I forgot how fun solving software puzzles can be!
It’s unfortunate our industry has become so boring. Because our society values capital above all else, quitting my job to do something fun would relegate me to economical loser status. We really need a shift to the left in politics and a way to redistribute the productivity gains this technology allows.
> AI is just a tool, if your ability to form relationships was easily displaced by a tool, then I'd say it says a lot about the ability rather than the tool.
In abstract, sure, but that is a much shallower discussion that many people are strongly disinterested in. It's generic, where a human decompile may not be (as a cloner I haven't looked into those projects). I don't like the idea of a group of fans of art frames saying don't bother looking at what's in the art frame, our conversation is equal.
> hit with a flood of low-effort ports
> hit with a flood of half-decent decomps
> It kinda just... totally eliminates that as a hobby entirely. I know people are going to be upset about this.
A rising tide lifts all ships, or however that saying goes. And a ship with a hole in its hull will sink to the bottom of the ocean. So don't sweat the surge of low-effort projects...
Your hobby is not going away. In fact, it's totally popping off right now. This is the perfect opportunity for you to distinguish yourself and lead by a positive example (whether that includes using AI tools responsibly, or not at all, is up to you). Just keep on hacking the good hack.
“Congratulations, your small hobby subreddit has been chosen as a default subreddit! Prepare to reap the rewards of 1000x the amount of:
- traffic.
- low efforts posts.
- users who won’t read or respect the rules and culture.
But the point here is that the rules and culture as they were no longer matter. They were there to ensure a level of quality that is now reached by an individual with a 200$ Claude subscription, so they should be discarded and the community will have to converge on some kind of new normal.
I'll have to take a look at this, but I've found that the top models do quite well at this work without any extra work.
The Windows Remote Desktop client has two bugs that have been driving me crazy for the better part of a decade. One day I got fed up and fed the binary to Claude, asking it to fix the two bugs. And it just did it. Patched one with some NOPs and adjusted a stack offset for the other. Produced a credible explanation for both, and in fact the fix worked. It helps that I was able to describe the bugs clearly, but still, it had to figure out the right spot among megabytes of executable code.
This is crazy, solving a bug from the executable is orders of magnitude harder than solving it from source code, this means that even with all the money they have, Microsoft didn’t bother to fix their shit.
It's harder in the exact way that AI is good at, though. You need to slog through very large amounts of boilerplate until you find what you're looking for.
It's not really more complex once you're practiced at it, just tedious.
Eh, there's a huge difference between having AI fix a bug and fixing a bug.
The latter implies that you've understood where and how it happens as well as noticed any other areas that may be impacted by fixing the bug or by not fixing it. AI just fixes it without thinking about anything else - which is great, but it explains why serious people (and Microsoft) aren't doing just that. We've spent decades worrying about quality, reliability and performance. Using AI is saying all of that was poppycock and it's faster to rebuild than to understand and fix. It's a different school of thought, and hopefully one that bursts down in flame in a few months.
You might have noticed throwing billions around has diminishing returns
That's literally true
That's absolutely false
Much harder than fixing from source, but much much much easier than making a bug report that Microsoft / Apple / etc will care about. You get better results from a wishing well than Apple Radar.
I am amazed by the jump in usability of models this year. Some little thing annoying you about some software, some feature missing just point AI at it and burn some tokens. (ok, maybe a few rounds of telling the model to do better are still needed from time to time)
vmconnect.exe (Hyper-V enhanced session) not going fullscreen after login when the window is fullscreen and you have to minimize and maximize to get fullscreen? Have your agent debug the issue and come up with an elaborate hook system that leaves the Microsoft files untouched.
Codex/Chatgpt electron app coming with annoying or missing features (Mini, Invite friends, no mcp hotreload, windows updates failing, ...)? Just point codex at codex and have it build an mcp Ouroboros to improve itself and enable dynamic js userscripts.
Deskflow clipboard sharing between devices missing file sharing, just have sessions on Mac, Linux and Windows collude to cook up a solution.
Yet writing proper documentation and reports that are human digestible still seems a bit out of reach, why would humans need to read anyway...
That’s unbelievable. To be able to locate and fix bugs by patching the binary.
Studying and patching binaries has always been done. This is not new.
The novelty isn't binary patching itself. Its the the barrier to entry. Manually locating an bug in compiled code and calculating offsets used to require good skills and knowledge in the domain. Here, someone described the symptoms in a prompt (reminder, still plain English) and got an binary patch without touching a disassembler. To me, that's unbelievable.
What were the bugs?
That would make for an amazing blog post. ou got one per chance to follow?
I dunno, this isn't surprising at all now. The latest models are just really really good at reverse engineering. I asked Astra whether it was possible to do a task with a binary (commercial) library. I didn't even ask it to reverse engineer it but it happily went ahead and decompiled it, found some undocumented functions, figured out how they worked and then wrote an example usage for me.
My thought: Claude may not have fixed the bugs, but avoided the problem for the poster’s use cases. “Fixing a bug”, in my mind, means solving all problems for all use cases. And Claude didn’t have to do that.
The insertion of NOPs is a great way to remove a function call that might have crashed the program. If you didn’t need anything that function did, you’ve solved your problem but not fixed the bug.
By that definition scarcely any bug gets fixed ever. Perfect is the mortal enemy of good.
Also I don't think it's an honest take. You took this stance because the change was done by AI. If a person patched MS app he uses that had a bug that bothered him for a decade you would be more appreciative.
I wonder if this is why I've seen a lot videos popping up on my youtube feed related to vibe coded clones of various commercial apps in the past few days. Everything from clones of flagship products from Adob to Microsoft Office.
Adobe product clones like Photoshop and Illustrator: https://www.youtube.com/watch?v=eFB79TYI-Vw
Adobe after effects clone: https://www.youtube.com/watch?v=5mi_tYSdkWQ
MS Office suite clone: https://www.youtube.com/watch?v=U_jTYMOlXio
No, we explicitly do not use reverse engineering. We do not decompile binaries, we do everything 100% clean room:
https://github.com/storytold/photocraft (inspired by Photoshop)
https://github.com/storytold/wordcraft (inspired by Word)
https://github.com/storytold/pdfcraft (one of the more mature apps)
https://github.com/storytold/vectorcraft (another app close to 1:1 parity)
(etc.)
Using REA for something as high profile as what we're doing is likely to result in lawsuits. We're doing everything we can by the books.
We cannot look at Adobe sources. Use of Ghidra is disallowed.
REA is probably great for personal apps and for abandonware, but I think if you publish the results and it's found to have decompiled the original proprietary sources in discovery, you might be in for a bad time.
Using an LLM is not a clean room, imo. Its just IP laundering. Which is fine I guess if everyone is doing it, including the companies you're stealing from. I just dont know what the implications will be for progress.
Licensing/copyrighting encouraged people to think up of new things, and new ways of doing something. Now we're just all copying eachother.
What intellectual property is being laundered here if it's not even the same ecosystem (C++ vs Rust)? If it's so easy, and you just need to rebatch something existing, why has nobody done this over the last 20 years?
Adobe almost certainly has patents covering aspects of Photoshops design and tools?
Microsoft famously has patents covering aspects of the ribbon interface in Microsoft office, which of course this suite must implement. https://en.wikipedia.org/wiki/Ribbon_(user_interface)#Patent...
Meanwhile, Wikipedia has a policy that says screenshots of applications have to be "as small a version as possible" in order to meet the fair use exception for presentation of copyright work. I can only imagine an exact clone of the interface could be similarly ruled to infringe on Adobe's intellectual property.
Copyright is so flawed... This would mean an hommage is a breach of intellectual property in general, even if you paint it completely differently (say watercolor vs oil)
I agree. The vast amount of data ingested and internalized by LLMs has effectively been "laundered". But they are so powerful and evolving so fast that no one can be spared of their impact. We have to to learn to live with it. Traditional proprietary software being "laundered" is just one part of the broader story...
all human culture and creation is a form of "ip laundering" i guess
Doesn't clean room typically apply to cases where the "dirty" team has legitimate access to copyrighted code, like the IBM PC BIOS which was published in the technical reference manual, and uses this access to write a functional specification for the "clean" team?
I'm not sure how clean room would apply to commercial applications distributed in binary form, as there's no way to look at even disassembled code without violating the license agreement and therefore being in breach of contract and subject to potential copyright infringement claims for copying or even continuing to use the software, let alone cloning it, and surely you're not going to be subject to a copyright claim based on familiarity with the application from merely using it.
The issue is that LLMs very likely have access to the source code of these products as part of their training data, and are then being used to generate clones. There's no barrier in the middle to ensure copyright violations don't leak
The reason why 'reverse engineering' has gotten so good is because what we're actually seeing is fully automated luxury plagiarism
It is very unlikely that LLM had access to original source code of Photoshop.
> It is very unlikely that LLM had access to original source code of Photoshop.
It doesn’t matter. It’s likely had access to the knowledge of somebody who worked on that source code.
I have worked at companies that have done clean-room implementations. I was not permitted to look at their source or interact with those teams, because I had been exposed to the source of what they were re-implementing in a prior job.
As a human with a mushy brain I was never going to remember the source, but the fact I had been proximate to it was enough to lock me out. You have no idea what the LLM trained on, so you can’t prove a LLM derived clone is a “clean-room” implementation.
It could have access for code of an older version with how there has been leaks of sources of various commercial applications over the years.
They have had access to plenty of A/GPL software. They are all highly problematic from an ethical and legal point of view.
If adobe use any AI tools to help develop photoshop, its all been loaded in. Or if its ever ended up anywhere on the internet, even in private
The first part (re: AI tools) is ostensibly not true. It all depends on the degree to which the data that they "don't train on" gets laundered and anonymized to the point where they are able to justify training on it without it being "yours" any more.
And I don't think that is publicly known at the current time. (If anyone has any tangible info on this, the please let me know...)
The practical uses are more applicable to laundering open source code covered by copyleft licensing like GNU.
This is novel Rust/egui code that I imagine looks nothing like Adobe code. I've never seen their code, but it must be a mess of old C++, right?
It's considerably faster than their apps (at startup) too.
Well I hope they trust their LLMs completely, because I'm sure Adobe is currently going over their source code with a very fine comb to build a case. And they have LLMs to help with comparing too.
People made new and better things just like companies without copyright too. In recent times companies made things worse too.
Please don't take offence by this, but:
If AI models can generate designs faster and produce work that is "good enough", what is the actual future of design tools and the design profession in your opinion?
I've tested this myself with some frontier models and the results are kinda impressive enough that it raised the question if design skills are already obsolete. If that is the case, what use are these tools now?
Good enough is not really good enough.
I've seen a lot of "really cool unique designs" that are obviously just a digested regurgitation of amalgamated corporate slop. Anyone wanting an actual unique design language needs to hire actual designers.
Those designers may use AI tools but the tooling isn't to where they can be completely replaced. I challenge anyone to show me an e2e LLM design toolkit that can actually replace a designer, not just one shotted "wow that looks so cool" character designs.
How do you know llms you are using were not trained on decompiled apps?
why would that matter? That sounds like a problem between Adobe and the AI companies.
Would love to see an InDesign clone that supports writing to and from their proprietary format.
The death of IP in the west. Good thing? Bad thing? Who knows. Definitely the start of something big though.
Mightn't be so bad if all the value wasn't being captured by trillion dollar corps, despite them benefiting off the labor and creativity of countless non consenting individuals
Those trillion dollar corps also employ millions of people in very high paying jobs. And the extraordinary stock returns have made their engineers (and countless public investors) wealthy across the past half century.
People wishing for the death of the mega corps, are directly wishing for the obliteration of millions of great jobs.
How many of those jobs do you expect them to retain as they build increasingly capable models?
How much of the profit will be redistributed to all the artists, programmers, scientists, etc. whose works are the foundation of their product?
The poorest 50% of global citizens hold the same combined wealth as the top 10 individuals.
How are massive and increasingly concentrated corporate profits going to help them? (Or anyone else for that matter).
I’m not sure people are “wishing for the death” in the literal sense.
We just need to apply the antitrust law and break these trillion-dollar corps into smaller competing corps.
There’s no reason for a single company to provide your email, browser, search, operating system, entertainment, AI, etc.
>People wishing for the death of the mega corps, are directly wishing for the obliteration of millions of great jobs.
Yeah, I'll take the obliteration of a few million overpaid American jobs over the destruction of hundreds of millions worldwide, the social upheaval that goes with it, the destroyed families and the direct and indirect deaths. Fuck your promo packet and fuck your cozy desk job.
Cry me a river
IP is not dying. It’s being consolidated to just a few mega corps. Good thing? Bad thing? Definitely the start of something big though.
Imaginary Property was always an illusion, one which is now crumbling quickly despite a lot of desperate attempts at maintaining it.
Everything was, is, and always will be a derivative work.
Wow. Maybe AI can win artists over after all.
great work, congratz.
The AI labs already figured it out:
1) reverse engineer the code 2) train a model on the code 3) use the model to write the clean code
Step 2 is the key “cleaning” process
So maybe a good strategy would be to use something like REA, put it on GitHub, wait for the LLMs to train on it, then just use the frontier models
/s
I prefer the word *laundering
LLM's have already trained on similar enough code
Even better, 1 and 2 done, just move to 3 and profit
Probably not, there's relatively little secret sauce to something like Photoshop. It's just a lot of grungy work that, I guess, you can now delegate to an agent if you have enough money and time.
As an aside, I've heard a lot of hot takes about how this is the end of Adobe, but I'm pretty sure it misses the point. The main reason people pay Adobe is because it's a familiar line of stable, well-supported, interoperable, and actively-developed products. There's already plenty of cheaper or free alternatives (Davinci Resolve for video, Capture One / Darktable for raw, Affinity for photo editing and vector drawing, etc), and if Adobe survived that, I sincerely doubt they're going to lose pro customers to a vibecoded app where half the stuff is probably subtly broken or left as a TODO, and that will be abandoned in a week, because the whole point was to get that 1M YouTube views.
Not really. Someone using Adobe Suite of products cannot easily switch to Davinci Resolve or Dark Table.
These are NOT the same thing whereas this one particular suite of clone products is exactly menu by menu, keyboard shortcut by keyboard shortcut and feature by feature a pure clone of corresponding products to that extent that they even have a separate dashboard just to track feature parity: https://github.com/storytold/craft-repo-status-app
It is NOT there yet of course but it won't be there in a year? That's an absurd claim to make. I have checked the repos multiple times and every new release keeps fixing something that wasn't working before.
At some point Adobe will go after the clones, copyright their interfaces, close standards, patent stuff etc. The barrier for software used to be implementation effort, but now that it's out of the window, it will be more like other forms of IP (music, physical products, etc).
> that will be abandoned in a week, because the whole point was to get that 1M YouTube views.
Longbets 1 week, haha.
We're working our asses off on this.
Most of the team are artists who use these tools actively and we want the replacements for ourselves. I'm a filmmaker, so you can imagine my frustration of being bitten by the "unsubscribe fee".
> being bitten by the "unsubscribe fee"
I’m curious how much is your LLM bill. I’m assuming the idea is that LLMs will maintain the software? I’d be surprised, given the token economy, that you’d have to pay less than those subscriptions to LLM providers.
great question
> I'm a filmmaker, so you can imagine my frustration of being bitten by the "unsubscribe fee"
What's the 'unsubscribe fee'? Are you referring to choosing something like an annual plan instead of 'Monthly' and only getting 50% back if you cancel partway through?
As you said you're a filmmaker, if someone chooses to obtain the rights to your film for a longer period instead of a shorter period and then changes their mind halfway through, do you similarly give them 50% back?
From what I can see, Adobe prominently and clearly displays from the first public checkout page the differences between 'monthly' and 'annual' options?
> From what I can see, Adobe prominently and clearly displays from the first public checkout page the differences between 'monthly' and 'annual' options?
It does now, because the FTC sued in 2024 and Adobe then corrected this, while settling in March 2026
You're doing God's work. Keep it up! PhotoCraft is really cool, and the pace of stabilizing progress is incredible! Don't listen to people who dismiss your work without getting their hands on it.
Have you guys done any talks or written any articles on how you are developing this software?
Scaling up to this kind of complexity even with AI is still a challenge, at least for me.
I am definitely interested in alternatives, but I am not interested in an alternative that isn't supported and maintained. Though I suppose at some point I can ask my own agents to support/maintain it?
No ethical qualms here but you guys should save your money and just use Gimp! It's already free, really capable, and comes from a good and pure place of making software for its own sake. That's the kind of thing that gets you users for decade, beyond any egui rust doo dads.
They are covering the entire Adobe suite, not just Photoshop.
I've been really inspired by it. Ever since that sort of joking project came out, "Malus," that would clean room reimplement GPL software under more permissive licensing, I've been chewing on doing the reverse for proprietary software. I've been calling it the Manfred Macx theory of disruption, for the character in Accelerando that would patent and then release to public domain profitable ideas before a megacorp could lock them down.
I've got my wedding coming up so haven't had the time I want to dedicate seriously to finding some proprietary things to disrupt, but I've had fun using this as an opportunity to play with subagent orchestration, open weight models, various harnesses, and local models, to see what happens if I let some LLMs churn on it. That resulted in a hodge podge researched list of potential targets: https://github.com/508-dev/genairosity/issues
One thing that stands out is a lot of software kinda does already have a FOSS replacement, it's just not really how people want it to be. GIMP being the representative example. Photoshop people just don't like it, I'm sure for not entirely invalid grievances. However the maintainers of these kinds of tools are often strongly opposed to LLM involved contributions, again , often for not entirely invalid reasons on these incredibly complex projects.
So the long and short of it is that as powerful as LLMs are, we aren't quite where some of the doomsayers are saying we are, insomuch as proprietary software is dead. Even with reverse engineering we aren't there. There's genuine labor, time, and expertise moats around most of these programs.
Personally I'm shifting to trying to find niche abandoned software with no export flow. Even if I can only help out a couple hundred people, it sounds like a decent use of my time.
I built a from scratch rust RAW processing engine similar to the brains behind photoshop and lightroom - its much more powerful and its not even complete yet
yep. it is unexpectedly everywhere. like, overnight.
LLMs are excellent at generating something that looks good, not so great at generating something that actually is good. All of those look really slick, but apparently the functionality is all broken
lol :), you're wrong.
I see this liquid software being the unavoidable future. A computer that just does stuff in whatever way you guide it, in whatever way you like guiding it. The models will keep getting better and faster up to the point where everything is just happening in real time, no OS no drivers no programs, just an entity that can listen to you and can move around bits to accomplish whatever you are trying to do. Wanna post on HN? Any way that you can code such an action, the computer can just manifest it for you, on the fly, however you want it.
What a dystopian future where your computer use will be basically “metered” by the token. At least right now we don’t have to pay per mouse clicks.
Local inference solves that.
We used to pay for phone calls not long ago. Now a video call is free.
Do you believe the average person can afford the 1TB+ VRAM required to run a competent model?
The expectation is of course that that will not be required. Competent models will become smaller and the compute to run them will become cheaper, once the market stabilizes. That will take some time of course.
Not today. But just like the average person can now afford a smartphone which is more powerful than a workstation not long ago, we will have enough computing power for good enough local inference, at an accessible price.
We're not in the 2010s anymore, consumer hardware is actively regressing. Smartphones are actually a great example to bring up, because I believe 2026 was the first year in which smartphone hardware did not advance at all (and arguably declined) relative to price, due to AI-related shortages.
Nobody can predict the future, of course, but I personally believe this pattern will hold. A decade from now, the idea of a consumer being able to purchase a personal device with more than 8GB of RAM will be a thing of the past, all significant compute will be done in the cloud using the massive amounts of hardware being hoarded by corporations as we speak.
I love such work. I hope to roll it into a package that anyone can use in the future. This work is only going to get more important because frontier models getting locked down will make this a lot harder over time.
For example, I use Claude as a bouncing wall for my thoughts and I pointed out that,
This was rejected for "Safety," Note, the image here was the Library of Congress' page on DMCA exceptions.Fundamentally, the idea that you can't reverse engineer things, make things, learn about biology or physics without permission is strange to me. These machines have been trained on the sum intellectual output of humanity, the global intellectual commons, and are being used to close off that commons?
I would be OK with their right to create such restrictions if they weren't lobbying the Government to restrict others, thereby ensuring that they control humanity's intellectual commons well into the future.
Perhaps I'm naive, but I think it's better for humans and the machines if we can all think, learn and build. But then again, I'm the kind of person who rejects the doomer pill.
Yay, all software will be open…Except the models.
use open models and abliterate them
I wanted to see if CVP approval changed this response, but it appears that with the release of Opus 5.5, Anthropic silently dropped me from the program, and has some strict new criteria in place to apply again, such as being credited for a CVE! I was only approved last month, too -- sad!
I have CVP with the new program (including mythos access) and still get constant denials for silly situations. Most recently I fed a URL to my agent from a security blog and asked if to add it to my obsidian vault with appropriate tags-- cyber flagged. You're not missing much. OpenAI and/or most Chinese models are much more lax in their restrictions.
This kills the talent pipeline, and it'll create a spam problem for the other folks because now people will try to github PR spam their way to getting on a CVE.
It's worth talking about the fact that you can't even talk about DMCA to a model trained on the Library of Congress unless you're one of the approved people. And that's before reverse engineering something or writing code.
So in this future, it sucks to be you if you're someone trying to make your small app more secure, someone trying to upskill, a tinkerer trying to bypass corporate lockdowns for a device they own (a recognized DMCA exception, btw), a teenager trying to learn about security...
It locks away much of the richness that produced hacker culture behind glass. You can look at their press announcements and PR pieces, but you can't touch.
And as they're lobbying the government for "sensible regulation," this inevitably leads to a future where computing is controlled.
It's the direction their existing reports are taking. They recently released one in September that talked about how they stopped "bioweapons." What were said bioweapons efforts? Oh, it was scientists using Claude for grant writing, paperwork and grammar. At national labs.
These people are basically proud of impeding real research to make better painkillers and study a neglected tropical disease, https://news.ycombinator.com/item?id=49651727
And this is being used to lobby against "dangerous" open-weight models because gasp a scientist might use them to write a grant! To make better antidepressants.
At what point do they start reporting someone taking apart an iPhone and trying to DIY a repair with a schematic as a thwarted "cyber security incident?"
A funny, but slightly chilling safety violation I once got was ChatGPT being unwilling to recite the full text of Article I Section 2 of the US Constitution, aborting as soon as it hit the passage about "three fifths of all other persons".
Another funny one was Claude's refusal to provide the original untranslated text of a passage from Dante's Inferno on copyright grounds, though in this case pointing out that no 14th century literature was subject to copyright anywhere in the world was sufficient to override its objection.
Agree on the talent pipeline. It can take a long time for someone to obtain a CVE that has their name on it. You don't start being a security researcher only once that happens.
Several years back, I was working on generating AVB2 hashes on top of modified Android distributions, to increase the security after an owner has made their desired changes. I was doing this before the age of LLMs. Among other things, this would've enabled the secure features to work again, and potentially reduce the risk of root access being usable by malware. But apparently I'm not a security researcher because I didn't get a CVE about it.
It's good that Stallman is still alive to experience this world. I think he never would've thought that freedom in software could come in this form, and as a long-time reverser myself, I've always held the opinion that his fixation on source code and the free software movement was not as liberating as it could've been.
"Source code? We don't need no stinkin' source code!"
Pretty sure Stallman would have some choice words to say about his idea of "software freedom" being twisted to include "relying on a corporation's subscription service for all your software needs and being unable to do anything yourself".
Give it a few months and open-weight self-hosted models will be able to do similar reverse engineering work. It is an extraordinary future we live in.
Please explain how this is more liberating than access to source code? Wouldn't this be reliant on continued and cheap access to the models? Even the open weight ones are heavily funded and not out of the kindness of their hearts. Just a few comments down is a thread about being blocked by the usual guardrails from the main providers.
But this is access to source code. All source code is now essentially open. If not now, it will be in a few months. You aren't any more dependent on the models then before , you just have the source code of everything that compiles on your computer. What happens in the future is anyone's guess, but for all software published until this point in history, closed source no longer exists.
It's "access" to "source code" as long as you can pay Anthropic/OpenAI and they don't raise their prices or lock you out.
No I mean sure, if you want to decompile it yourself from scratch, but once anyone does, then the source code is available to anyone. So from that point onwards, no one is dependent on LLMs to decompile it again.
GLM 5.3, an open model, is also quite good at reverse engineering.
really creeped out by the ai bro takes, specifically the ones "at least in a few months", almost as if there's no engineering principles anymore and just hype
Eh, decompiling is happening now, at current LLM capabilities. The resulting codebase compiles into something that is functionally identical to the original. What engineering principles exactly do you need beyond "works exactly the same as the original product", in this particular context?
I might be missing something but... How exactly is this better than telling Claude for example to "install and set up a full RE environment including Ghidra" on my local system and get to work? Like what does this do that my current RE methodology doesn't?
Probably it is not better. I think this is aimed at people who do not have a “current RE methodology”, do not know enough to specify things like Ghidra, etc., but who do have a desire to feel like they reverse-engineered and can reliably predict that a conversation with a chatbot will make them feel that way.
Vibe-reverse-engineering
Reverse Vibeneering
I mean thats not why I have Claude decompile things. I have it so it because my software should work exactly how I want it to.
Everybody has their own set of skills and specific scripts and tools to do this stuff. You might use Ghidra as the the kernel of those workflows, but you still want something more than just Claude freestyling, at least for now.
(Who knows if this'll be true 6 months from now.)
I don't have reverse engineering experience, though I have all the prereqs to learn it. Anecdotally a few days ago I told my slow local Qwen3.8 in Pi harness to use Ghidra CLI to decompile a certain executable and it got entirely lost. Today I was linked to this, have it chugging along now, and it's making some sort of progress towards unpacking this thing. I don't know if it will nail it this run but it's a real improvement; I definitely have some useful info I could bring to another prompt or for myself if I cared to try manually.
But I just found an even easier way I should have thought of first - someone already dropped a reimplementation a couple weeks ago.
Why not just, like, ask Claude about it?
"Hey Claude I want to create a fanmade Game Boy game, give me recommendation of tools, libraries and workflows for it" and boom.
You can use AI to learn stuff, instead of just using it as a Pokemon.
I don't know either. My only guess is that this has some helpful context for less powerful models, maybe.
We're kinda far into this LLM thing, maybe it's time to start selling harnesses by leading with how some examples were solved faster with this and how it saved tokens, or similar?
Because it's already been done for you? Sure, you can spend your own tokens on it, but it's probably cheaper and more time-effective to use something that already exists and does the job.
I don't think it is better than just doing it yourself in Claude Code (in fact worse) but some people like cute UI for everything I guess.
i just tell it to use radare2
We're using vanilla Claude Code for ArtCraft apps [1], but we are especially careful not to touch Ghidra. We don't want decompilations or reverse engineering to spoil the work we're doing and expose us to copyright infringement.
[1] https://github.com/storytold
Slightly off topic. I looked at REA's Android reversing support and found it still uses jadx mcp. That kills large scale APK reversing. jadx takes tens of minutes to preprocess an APK, build a code relationship database, etc. Even headless mcp is no exception. I basically can't use it to analyze APKs at scale, like 100 large commercial APKs in a pipeline, or 100 preinstalled APKs from a phone ROM to hunt bugs.
So I built droidasc. No memory bloat, no parsing slowdown. Analysis is in milliseconds. Global xref on a 300MB APK takes 1.5 seconds. I used it with codex to analyze phones from 3 different brands and found 2 RCEs and 5 root bugs in a few days. I plan to detail these at Black Hat Asia 2027. A friend used droidasc to scan various bug bounty targets at scale and found 10+ RCEs. Way, way faster than jadx.
https://github.com/MG1937/ASC
Interesting, might be fun to try my hand at it sometime. Is the best way to actually get the APKs to just download them from an apk mirror site.
Yes! APKMirror is a good choice. AndroZoo works too, but it's academic and you need to apply for API access.
APKPure is pretty good too, they don't have captchas for downloads so you can make a simple script to pull the APK from them based on the package name.
You should be able to get free APKs straight from Google via Aurora Store.
I’m surprised people don’t get refusals running this, or are they running it with cyber-enabled models like daybreak?
In my experience, even just hinting at reverse engineering to models from Anthropic or OpenAI leaves them extremely sensitive to refusals. After all, the same techniques used here can be used to find exploits.
Or are people using it with local models?
Curious.
I handed Astra Ghidra with some printer driver exes/dlls loaded, and a Windows 2000 VM with the software installed and said "reverse engineer this printer driver, it's hooked up and powered on on /dev/ttyUSB0, tell me when you're ready to print a test page". No refusals.
I did more or less the same and was insta-banned by OpenAI for life. Stay safe out there.
What? What software did you try to reverse, Ive never heard of anything even close to this
Really just a few attempts at cybersecurity related things (not even RE) with ChatGPT/Codex triggered a warning of breaking ToS by email. A few days later I was banned. No appeal.
My conclusion is that OpenAI will help you with cybersecurity and never really block or prevent you from doing the work, but then suddenly warn and/or ban you. Anthropic does the opposite, where they trigger safeguards all the time while working, but never really ban you.
Reverse engineering and decompilation is not currently blocked by the models, leading to the proliferation of LLM-assisted decomps/recomps. If you ask it to find bugs you're more likely to get a refusal, but simply pointing an LLM at a Ghidra MCP server and telling it to trace program flow, rename functions, or answer questions behaves like normal.
Yes, it just rejected me once with Claude Code Opus 5.5, but it works fine with Codex 6.1-Sol (in my experience, it was the other way around)
Opus 5.5 hasn't given me much trouble with games, but for DRM and licensing checks I usually have to drop down to 4.8 or 4.6.
I think people said that Claude refuses to work on RE sometimes.
I've had no issues with OpenAI.
Ironically, I've had the opposite experience. My OpenAI account was permanently banned for doing too much RE work, meanwhile the worst I get on Claude is a downgrade to Opus 4.8.
Oh, that's interesting. Mind sharing how much was 'too much'?
I’m not the person you’re replying to, but for me just a few attempts at doing cybersecurity related things (not even RE) with ChatGPT triggered a warning by email. A few days later I was banned. No appeal.
The Chinese labs aren’t constraining their models as much as Western labs do. That’s why HF used GLM. If you remember the early times with how Gemini (Bard then) and Claude refused almost everything, you’d despise the idea of guardrailing. (Its important in some areas though)
I'm using IDA Pro's MCP (Well worth of money in the past), it may hit the infamous artificial cyber wall at any time.
Try GLM-5.3, it worked pretty for me. It also worked well with Radare2 or binary ninja if you don't have the muscle memory for idapro.
This makes no sense to me. There are innumerable reasons to reverse engineer software that have nothing to do with security, some not only not prohibited, but in fact specifically authorized by law.
It'd make more sense to reject reverse engineering commercial software on the grounds that it likely violates the software's license agreement, and would therefore subject the user to potential breach of contract and copyright claims.
But this would also apply to uploading pretty much any non-self authored document to the LLM in the first place, albeit with fair use as a possible defense after the fact, so it still doesn't make much sense.
With C#, whenever codex need some info about the packages we use, it just use ILSpy to decompile the nuget instead of checking the docs. its much effective though. Even in my own packages, it use to gave back the bugs or gaps that i need to patch up.
Does anyone have insight into the sudden trend of game decomm / recomp and mashups like Minecraft in GTA, Mario in Skyrim, Cod in Life is Strange, Among us in Portal 2 and so on? What triggered it? Is there a specific new tool, technique or LLM that sparked this interest? I am curious.
Update, found this article to cover the timeline well. https://knowyourmeme.com/memes/cultures/ai-video-game-mergin...
This tweet is what opened the floodgates: https://nitter.app/chasmmmmmmmmmmm/status/210137541186725513...
Opus 5.5 being cheap and extremely capable + added YouTube virality.
I’ve been doing decompilation with LLMs for 4 months, and 5.5 is so insanely good I now can finish Windows 95-98 games in just a week. And by finish I mean get byte-to-byte exact output from your decompiled code to retail with clear names, semantics and no “Ghidra smell”. With older Opus models you would have to manually guide them and inspect ASM yourself otherwise they will just burn through your tokens (which also used to cost more!)
Example: https://github.com/sushi-shi/homm1-decomp
And I have way more on my gh.
People love a good crossover?
I have extremely mixed feelings about how this impacts things like game decompilation projects, but when it comes to cracking DRM my feelings are only positive. (Although I'm sure those on the other side of the DRM are unhappy)
When you crack some DRM, the DMCA makes it challenging to share your work with the world, regardless of the morality of your use case (say, repairing one's tractor). There is no longer a pressing need to share that work, when anyone can just say "computer, sync my spotify collection to my jellyfin instance, using correctly tagged FLACs", and it goes away and does it from first principles.
So you don't care about artists then? Those Spotify streams is what makes them income.
Good point, I'll add "sync the jellyfin playback stats back to Spotify" to the prompt. Maybe with a multiplier for artists I want to support extra.
I don't know how much this matters as we get into a world where any software system can be composed in a few hours.
As the value of a particular piece of software decreases, the value of reversing its specific implementation does as well.
Reverse engineering is about discovering specific methods or protocols. Not for porting entire implementations to new platforms or products. A lot of the specific methods and protocols have been sucked up into the LLM weights, so the need to go digging for these patterns is dramatically reduced.
The biggest application I see here is with digital archaeological work (the opposite of new things).
Oct 5 ~4.2k Oct 7 ~10.3k Oct 8 ~17–18k Oct 10 54.6k live
Going from ~4,226 to 54,614 stars in ~5 days is about +50k. Even stranger is forks: 438 → 10,376 forks. That is an unusually high fork acceleration. Current forks are about 19% of stars.
Reeks of botting, some comments here also look artificial.
The setup part made me chuckle.
The copy this prompt into your agent seems like an evolution of `wget X | bash`-style in all the worst ways.
I wish in few years local LLMs/AI be more available and powerful so I could decompile and extend childhood games. Tools, missions... Damn. My dreams.
I’ve been trying to rebuild and modernise an old abandoned ms dos game, simply by pointing Codex (6.1 Sol) at the game directory, and it works insanely well. It’s able to understand the data formats, unpack graphics and sound assets, and reconstruct game logic.
Hmm, probably not a good idea if you want to build apps publicly. But, the most interesting about this project for me is how peoples reacted to it.
Morluto is a legend.
Started in the repoprompt (https://repoprompt.com/) community.
Good stuff.
I tried but can't figure out what's this tool's purpose?
> Copy this into your coding agent:
> Install REA and connect it to this coding agent using npx rea-agents@latest setup. Show me the setup plan for approval, then verify the installation.
We have achieved the next evolution of installation by `curl | bash`!
You're right it's much safer to click next > next > next > next > finish.
A big difference between the safety of "next > next > next > finish" and "curl | bash" is one of them is dynamically loaded from an external source that could change between runs, and the other can be fully downloaded and vetted in a single check, and then once it's safe, it's probably safe 10 years from now.
Every download of a piece of software could be unique.
Which is why we have sha1, md5 and sha256 hashes on display, so you can validate with a very high level of certainty, especially for sha256 at least for now, that the file is the same one.
We have existing paradigms for this.
Additionally, installers are signed with certificates on Windows.
All of these are strictly more trustworthy than curl | bashing.
Yeah installation has always been such a security issue. So many programs are just random links that download a file. You have to trust that the host has not been compromised all packages that were used to build it were not compromised etc.
With ai models getting better we may be able to do analysis on the actual underlying bytes of the files we download to properly scan them for malicious code patterns and build systems which sandbox programs and watch inbound and outbound traffic/ system level actions from them and flag suspicious requests for further analysis by smarter models.
REA shows that ai are very good at understanding low level code and reverse engineering it so this could potentially be applied to application level security aswell.
Or we could just use Nix.
I mean to be fair right under it they give you the command if you want to run it on your own. But yeah
Maybe we should bring Xara Xtreme for Linux back?
Anyone used this for firmware files?
Thanks for sharing and making this open source. Starred and followed
What makes this better than just giving an agent radare2?
My agent usually goes and installs capstone itself.
>installs capstone
hello fellow Claude user
Oh I'll be trying this out for the OpenBFME project!
At some point I think it would be better to just redo the thing from zero like was done with Beyond All Reason.
The mechanics of BFME are the important part, not the intellectual property of the films. The way the units move, the resource system and power points are what make it special. Gandalf and Lurtz are simply wallpaper on top of something that is already very competent.
the chickens are coming home to roost
So what do companies do about this? If every piece of software can be reverse engineered are all companies just going to move to saas only?
Are we getting to the point where every API can be reverse engineered too? If so, that's not safe either.
Ooh it can recreate old games from executables, I wondering if it can do Motorola MC68010, I want to play Stun Runner in 4K.
https://www.youtube.com/watch?v=tByxdDiRdPM
Now there are no excuses to reverse engineer the most notorious closed source binaries out there including from Nintendo's system software to CUDA, and nvcc from Nvidia and make it all "open source".
The only problem is the lawyers at all those companies will be readying their lawsuits, and given they have tons of money; they do not care and will come after anyone.
What do you plan to do with a reverse engineered nvcc?
Why wait months for Nvidia to fix their compiler bugs when its far more faster to do it in the open? Libraries that invoke nvcc don't want to wait either.
So really closed-source compilers are a hindrance and you're subject to unspecified time-frames to even get to fix your issue rather than doing it yourself.
Intel and AMD have already open up theirs without question. Nvidia is the one who claims to be supporting openness in AI but can't even open up nvcc let alone CUDA.
I think you will spend far longer trying to fix all the bugs introduced by your AI decompilation.
And then you can use Claude-Magnum to fix the issues.
And then you can use Claude-Oeuvre to fix the issues.
And then you can use GPT-Cosmos to fix the issues.
And then you can use Claude-Omnibus to fix the issues.
And then you can use Claude-Logos to fix the issues.
> CUDA
That's insanely meta.
I guess everything is going to become recursive. RSI on models. Everything feeding back into its own hill climbing optimization.
Humans nudging it up the hill further and further by using new geometric guidance.
How long before the backend of a bank, or an insurance company, or a hospital will be reverese engineered?
"Reverse engineering" the backends of banking systems is my full time occupation at the moment.
Think of it as a more persuasive means of obtaining documentation. Not as a nefarious act.
Do you have access to the backend of a bank?
I really doubt there's anything interesting in there, judging from experience of looking at implementations of mobile banking apps and websites. Horrendous API, some intrusive scanning/probing/fingerprinting code and usually pretty weird ass ad-hoc cryptography protocols + some mechanism to try to lock API use to OS vendor giving a blessing for its use, so usually Google has to say, "ok, you can call API of your bank to access your money" or whatever on every single use of the app. (which is one of their ways to increase moat around their walled OS gardens) Otherwise you're out of luck using a mobile app. Websites don't have this issue, yet. Mozilla doesn't gate my access to bank APIs.
Backend will have some shuffling of data around some ledgers + a lot of ceremony around auditing + shit ton of CRUD mess and arcane connections to other systems/institutions. Probably the nightmare of nightmares codebase, if frontends are any indication. :D
Same with healthcare ime. I'm more on the human services side, but I was reversing their apis since pre llm days when I was a relative novice. Many don't even minify, so you can step thru the frontend src in devtools.
And yes, if frontends are any indication, just seeing tip of the shit-berg
There are some days I just hit a raw debugger in AuthForRealThisTime() under a three-paragraph jsdoc written by bot which is itself under the commented-out Auth() function. So many questions arise, and I can burn hours rabbit-holing the stack