The quality is quite high, but an observation is the direction they're taking these models correlates heavily to the usage demand of China vs the West. Specifically, they're immensely focused on t2v for action / high effect shots. There's one human reference shot in the entire release page, and that one doesn't focus on dialog at all.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
It's not about cultural difference but simply what's easier.
Action sequences are forgiving. There's lots of motion and shiny vfx to fudge things over. Rapid cuts mean each shot can be short enough that the inevitable accumulation of AI hallucinations from frame to frame doesn't get distracting. Sets can be generic: if you're generating robots attacking New York, the viewer isn't going to have time to track if the bodega on the corner is in every shot, or if the Chrysler Building switches place around town. Shots can be practically from different cities and nobody will notice.
Carrying over an actor's performance is the opposite. You want long shots and impeccable scene stability. You can't just throw some additive-blended particles on the actor's face to distract from its generation deficiencies, like you can do for the Marvel scenes.
Bytedance is a corporation looking for clout, not to help filmmakers. They're releasing a model that makes them look good, and action sequences do that.
It's a counterintuitive thing to viewers who have long been told that vfx is the most expensive kind of moviemaking. With gen AI, these somewhat convincing Marvel pastiches are trivial to produce, but a sitcom episode is utterly impossible.
I hear this a lot, but I observed this is no longer that simple.
Like Green Book (2018) made over US$70.7 million in China.
And one of the most popular films in 2021 was Hi, Mom (2021), made over US$785 million.
Obsession (2026) is another one that is doing amazing in China, absolutely beating high action films like Supergirl ($12 million and going vs less than $1 million)
> Obsession (2026) is another one that is doing amazing in China
I haven’t seen the movie yet but I’ve been a subscriber to the YouTube channel of Curry and Cooper [1] for a long time. I’m looking forward to watching the movie.
I wonder how many of the other people that went to see Obsession were already fans of Curry and Cooper.
Supergirl is basically a lazy slob who is born with godlike powers, I can see why that doesn't play well in a culture that prioritises hard work and self-discipline.
I think it makes sense that historically American film focused on narrative based dramas and slow burn horror. Think of classic Hollywood hits such as Citizen Kane and Psycho. This is even apparent in the recent success of Obsession, which really follows strongly in the trend of Hitchcock, magical realism, human drama, psychological horror.
Compare this with historically successful Honk Kong films: Kung Fu Hustle, Police Story, Ip Man. I mean there are examples of Hong Kong films that are slow burn human dramas, like Chungking Express and Eat Drink Man Woman, and successful American action films like Diehard, but I think its clear that Hollywood film didn't start out with action movies, and Hong Kong film didn't start out with slow burn dramas, but adopted these genres afterwards; you could even claim that successful Hong Kong films like the Bruce Lee films of the 70s fueled the American appetite for action in the 80s, which was the most prominent era for that genre in the US.
As a corollary to this, it becomes clear that the Korean and Japanese film making industries are not as culturally distinct from Hollywood as Honk Kong film making is. For Korean film, the reason is obvious: Koreatown in Los Angeles is right next to Hollywood, in some respects is a part of Hollywood, and there are deep ties now between the Korean and American film making industries for that reason, and Korean-Americans have an overly high degree of representation in American media, and Korean productions often come to America to shoot (think of the end of Squid Games: it's almost certainly the case that they shot in LA because a producer has a cousin or something that works in Hollywood).
Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. But then I remember that I spend upwards of $10k on inference generating well over 50k images for storyboards, training models; and probably creating almost an hour of video (I assume). Yeah, I get that things can be economic if you don't use the latest models (ran some case studies on this), but the latest models are the most fun to work with. It doesn't scale as well as "vibe coding" stuff together on the weekend. And when things work really well its almost as if you're seeing an zoopraxiscope come to life for the first time; and you just want to keep going.
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
I went to Art Center for film. Loved it. But ended up writing software instead of shooting movies (while still also handling a lot of visual art direction, graphics work, UI, 3D animation, etc). Now I feel like we're starting to be roughly in the same boat as far as using prompts.
What bothers me is that every piece of content generated this way helps flood an already saturated market for content, while slowly degrading the expectations of what people see, to the point that no one will bother with shooting or animating anything anymore. Even if it's 50% worse, it's 90% cheaper, so the economics argur against producing any new physically made content. Simultaneously, it's cannibalizing all existing content. This points toward a feedback loop, like a snake eating its own tail. And even though Hollywood blockbusters have followed that pattern for a couple decades, it's demoralizing to me to see it enshrined as the future of film (or to hear from someone who makes films that it would be a preferred mode of creation).
MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.
Perhaps refined ComfyUI workflows will squeeze more quality out of it, but it's definitely not in the realm of Seedance, and Lightricks is training LTX 2.5/"LTX-Next".
Note: Those results are a little misleading the good LTX results were generated with a good workflow in ComfyUI, and a prompt expanding local model.
Whilst the H3 result was generated with the raw api using his raw prompt.
If you ask ChatGPT or other AI model to "improve" your prompt (with cinematic, good lighting) generally, you will get also very good result from H3 also.
My own H3 test show that H3 (API edition) is undoubtedly better than LTX in prompt adherence ! The real comparison will be with the edition of H3 we get to run locally.
Awesome! Have been a bit in the dark of the latest models, will have a look. I am due an upgrade for my local machine and GPU, so that excites me as well.
I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models.
It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)
Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.
Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.
Thanks for the input! And no I agree with your assessment of the website, I have some plans of a much simpler redesign soon, and I definitely value that critique. I think my "digital garden" has been through at least a few dozen of iterations, and sometimes I just get too carried away with it. The good part being, the next iteration always starts out better than the previous one (or at least I hope so :)).
There were tons of continuity errors, but I suppose bad film makers make those too and often get away with it. The worst was the gift that was boxed in one shot, then open and containing a gift that didn't fit in the box. Another more subtle one was after they bumped into each other they were standing in front of a wall suggesting she just ran out from the wall.
I dont normally like AI videos but this was just amazing, by next year the tech will be perfect, cheaper, we are going to AI videos everywhere whether anyone likes it or not.
people like it. Maybe not older generations, but we're about to get a fresh new generation of people that will not have a lifetime of experience looking at human generated content, and so have no real bias against AI content.
And then every generation after that will be born into increasingly a AI-video rich world.
The older generation loves them the most. Sometimes my mom sends me several AI generated reels in one day and she doesn't care if it's AI or not she just finds them funny.
Older generations absolutely love AI slop. If you haven't used Facebook recently, it's all old people sharing fake videos of animals doing silly things.
Well I find that after a certain age the behavior of old generations reverts to the same childlike patterns of young generations. It’s always the people in the middle who are getting screwed, they have to be the adults.
Right now we're going through multiple concurrent but slightly non-overlapping transitions where AI is almost good enough for something, and seeing a predictable surge of spam/bad quality crap people don't like, followed by a whiplash effect when it reaches human or superhuman levels of performance.
I think most developers would agree that the Cursor vibe-code era sentiment towards AI for coding was right in the sense that it wasn't that useful then, and a lot of people did and do stupid things with it, but it really did deliver quite substantially on its promise just a couple years after the "consensus" was often "that's never going to work".
It's very strange to me how consistently public opinion shifts on this problem. In the long run we're all dead but whether it takes one or five years, you probably will be there to live through it. I assume whoever is downvoting you is just thinking with their consumer-brain or "I don't want that" (the way it is now) rather than that this will be no different than video games, or wasn't a child who watched youtubepoop or other weird stuff because it was funny.
People are still figuring out the social norms surrounding "deepfakes" and I'm convinced there are some versions of the future where it becomes normalized for content under essentially fair use/free speech doctrine, and many where it's used to restrict free speech in the small number of jurisdictions where public figures would currently be allowed to be featured in this content.
Obviously there are a lot of ways it shouldn't be used but I want to live in the future where it's something we use to have fun, where reasonable, rather than a dangerous/sketchy taboo
UCLA is often used as a stand in for many colleges in movie sets (notably Harvard, as one of the buildings in UCLA has the Harvard crest on it specifically for this purpose) due to its proximity to Hollywood. It likely ended up in the training data that way
Yep, I notice it all the time in movies and tv shows. They were filming the movie "Old School" my freshman year there, I thought it was so cool. Makes sense it is in a lot of the training data
Seedance 2.5 looks amazing, but MiniMax H3 is going to be open weights within 24 hours: https://fal.ai/minimax-h3. According to the ComfyUI team, it should even work acceptably on mid-range consumer GPUs like the 3080.
I'd honestly take the slight quality hit for more control and lower costs.
This. I'm very much happy to see minimax-h3 be released. I have so many projects that I have on the back burner that I could complete within minutes instead of hours and hours. "Typography, UI, and Graphics".
New model day, here we go again! It’s a week of madness and all the wrong advice… and at the end of the week, the model is either dead or a massive hit. There is little middle ground.
I’m watching H3 to see if it fixes the nonsense that LTX introduced and what happens next… WAN is due, LTX says a new model is coming, Flux3 does video.
It went from zero video models to a lot of options due this year.
I don't think audio, image or video generation should exist. I haven't seen enough positive applications to justify the amount of harm these tools are being used to cause.
What are the positive applications? This all seems to be generating fake images/videos just because we can, not because it solves any real problems. It makes me wonder where the money is coming from to fund this stuff. Why would a regular person want to generate fake videos except for purposes of deception or debauchery?
The lack of creative understanding is astounding: all these generative media AI are fantastic for education applications. Sure, the vast majority want to use them to jerk off, but that's what humanity does anyway. Education is where these are extremely useful, and that is for the aid in conveying understanding to those that need the Jazz Hands of an image or audio or a video presentation to grasp the concept they are not.
While I generally think it's balanced against a lot of negative applications (misinformation, deepfakes, etc.), there are some good applications. I think regulation just hasn't sufficiently caught up with the harm they can cause.
To give a good use case (audio generation specifically), the 2002 game Morrowind has had a resurgence lately and it's a really great game, but it lacks fully voiced lines for dialogue. A community project called Voices of Vvardenfell uses ElevenAI to generate voice lines from the voiced lines that do exist in the game. They don't take away jobs from artists because it was a monumental (unpaid) undertaking from community volunteers to do this before (there was such a big project but it was abandoned last I saw).
So as with any emerging technologies, I think it just needs to be properly regulated.
Morrowind wasn't meant to be narrated, NPCs blur out walls of text with hyperlinks in them leading to more walls of text. If I had to sit idle while someone/something was reading all of this to me, I'd be bored out of my mind.
Some voices are “$politicalParty should win at all costs” and “my ex-spouse should be humiliated” and “my boss should go to jail for the pittance they pay me”, and instead of text opinions posted or voices shouted, there’s video gen to amplify them. They didn’t have the skill to make a convincing fake video themselves but if tools are at their fingertips… temptation.
Really not ideal although obviously amazing artists will do incredible things with the technology that many of us will enjoy.
It's going to be just like social media. A flood of content from the talentless masses exposing us to their lazy and shallow ideas they just cannot keep to themselves. Once in a while a diamond in the rough will be found, but at what cost?
The long tail dream didn’t really happen in any single market which was democratized the same way in the past. Why do you expect something else now, especially on an already oversaturated market?
Now they cause harm because there's people that may believe a fake video is real. But people is developing skepticism about what they see in videos and learning to think that they may be generated, before taking them as true.
It will be solved. Walleted, vertically integrated Apple can do a digital authenticity signature and guarantee that the video you see was at least shot on iPhone, and I won’t be surprised if it will happen soon. This + some lidar info will be enabled and everybody except minority will say that this is a great thing to have.
You've seen movies before. You know that they can make fake videos look realistic. We have an entire industry built around it, so why is it only a harm now?
Since it took a certain amount of time and effort, as well as skill in the past, it was a significantly lesser problem than it's becoming now. Especially with video footage.
In his study, "Nearly 2,000 witnesses can be wrong", Buckhout performed an experiment with 2,145 at-home viewers of a popular news broadcast. The television network played a 13-second clip of a mock robbery, produced by Buckhout. ... The people at home could call a number on their screen to report which suspect they believed was the perpetrator. The perpetrator was suspect number 2. ... Approximately equal contingents of participants chose suspects 1, 2, or 5, while the largest group of participants, about 25 percent, said they believed the perpetrator was not in the lineup. Even police precincts called in and reported the wrong man as the one they believed committed the crime.
If in the future they can create a reward model to aid with creating super human levels of appeal tailored to me, I would considered it one of the best technologies ever made. I'm not sentimental about the source, pretty videos and pictures make me happy
Fyi based on what people are saying on twitter, seedance 2.5 is ~2x as expensive. For a 30 second generation at dreamina it costs 1440 credits or around $15
I was trying to get at that its 2x more expensive for businesses, and personally I dont think 2.5 is worth double the variable cost. At least initially
Wait, but you compare the cost of animated show or a movie which is in its final, acceptable quality!
The cost of that $15 per 30sec is NOT FINAL - you'll pretty much end up deciding you want another take - either due to some inconsistencies, non-physical movements, or anything else, which screams "that's fake".
So the cost is most probably much higher; how much higher? We don't really know, as there are virtually no acceptable shows created by that tool yet (until there is something I don't know about).
>> The cost of that $15 per 30sec is NOT FINAL - you'll pretty much end up deciding you want another take - either due to some inconsistencies, non-physical movements, or anything else, which screams "that's fake".
Thats true for grok and other models. You need to generate endlessly and cut endlessly and stitch them up together.
but for seedance you pretty much get stellar output which is extremely consistent in a single shot. Like there is a very good chance you will like what you see.
If you want something vague and generic and don't particularly care about the details, then maybe you can get want you want on the first try, but you'll likely just get what everyone would call "slop".
If you have something very specific in mind, then expect to spend several dozen generations before you get what you want, or something close to it, and then you may still have to spend more time making final adjustments manually in an image editor.
...and this is just for one image. I have not tried AI video production yet but I imagine it would involve considerably more effort.
$15/30s is $1800/hr, so assuming 1/10 generations are good, that's $18000/hr for final content. A few orders of magnitude below professional movies for sure, but still not trivial.
It has improved a lot, but these demo reels still have all AI video issues. Flash cut salad (including the scenes that should have longer cuts), unnatural motion that looks animated, unprompted YouTube-face acting, etc. Admittedly it's all a lot less pronounced in this version.
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
This seems extraordinarily good to me, compared to what I've seen before. Their washing machine advert example seems like it's as good as anything else on social media. I'm shocked by the quality and the coherence they're able to maintain, I assume it's really good at using those reference images they mention in their prompts. I think the only one that's noticeably bad is the concert hall, where the first few seconds show almost empty stalls and then toward the end it shows a full audience, and that audience also looks a bit off.
The sample video looks incredible and the ability to keep details consistent for so long is impressive but it _still_ looks screams of AI in every single shot. I can’t put my finger on why, something about the way the rooms are put together, the expressions of the faces, the movements… It’s very unsettling.
I think we as humans are just really good at being able to tell. CGI has been around for decades and yet we can still look at it and say "yeah that was CGI". Uncanny valley I guess?
do you remember that article from google where it would hallucinate as if it was on acid/psychedelics and people going crazy saying this is the AI "imagining reality" and then shortly after some Googler went on a very brief media tour saying LLMs were sentient?
its crazy how much people extrapolate, the consensus back then was "cute but we won't get there" then suddenly we got GPT 3.0, 4o, 5.x, seedance
what people really underestimate is how much faster progress is now with AI
You know what, in my previous life as a filmmaker I could've only dreamed of such a thing. Filmmaking is an art form which you cannot do alone. Outside of a lot of time and money, you need cooperation of a number of people and each day of production you end up accumulating a set of compromises to your vision. Your taste is what makes you tolerate that or not, and it's exhausting. It matters if you're the author since at the end of the day it's your name on it, not the crew (as much).
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
I think for me it's just missing the "essence". Making a film is hard, but that labor with multiple people who are passionate about the project is just as important as the end result.
I've been trying to generate concept art for a game idea. I figure, I can't draw so why not have something else bring my vision to life. After dozens of failed attempts and burned tokens I decided to just hire someone. The experience was night and day. They provided thoughtful details to help tell a story in the image. They got to tell me what they added and why, etc. This isn't even their project but their passion for it showed in the end result.
These models are impressive, but I think people watch movies because of the essence as much as the movie itself.
Was just going to say... Even if all AI gives me is the opportunity to bring my visions to life even somewhat terribly... That's a lot better than I'd be able to do without it. Which is for them to go nowhere. I'll take it.
I am not a film maker but I had some sketches on my head that I always wanted to try to make. Even Gemini scratched the itch for me, and my folder “projects” is growing with other ideas that I’d like to sketch
The thing is, it has the same essential property that you don't control it any more. The problem with these tools is the lack of fine grained control means there is no room for you to express your individual creative input. Your total input is a few sentences of a prompt, then the AI did all the creative part. If the creative input is what you enjoyed, it is actually not that much more (or even less) here than it was with traditional film.
It will change soon. Ai is enabler. Again, as with photography and especially digital photography- everybody can click button but not everybody is photographer. Ai helps to raise the floor, not the ceiling
This is a matter of effort and direction and not an inherent flaw of the tools. Look at Apple's recent image tools to change perspective. Tools to improve artistry are improving.
Another aspect is that if you struggle and struggle and create something great, people can see the work you put in, and you'll visibly stand above the rest who don't have the same level of perseverance. But if anyone can create big studio quality visuals with a prompt, how do you stand out? "When everyone's super, no-one will be."
I often see this kind of passive-aggressive characterization of the process with AI that's meant to bait people who obviously are expressing their creativity with these models.
I hope no one will fall for it: even when you emphatically reply "I spent the last year building the pipelines I use" it'll still be treated as equivalent a couple of sentences relative to the old ways.
-
At the end of the day, the painful lesson some people are (re)-learning is that creative expression doesn't have to be tied to any specific process. A lot of people fall in love with creative expression + process... but some fall in love with just the process, and some fall in love with creative regardless of the process.
AI does not favor those of us who loved the process (that's me with coding), but it's definitely capable of allowing new people to experience what it's like to have creative input in your mind and see it expressed outside of your mind.
Often in making art you need to love the process. Otherwise you’ll never get good enough. Most filmmakers don’t even watch their films after they have spent years with it discussing, planning, shooting, editing etc. There is lots of travelling, new friendships, comradery involved and in a way it becomes the whole scaffolding of your life. Replacing this with prompting is probably not interesting to this particular group of people, but might be appealing to completely new set of people.
There are people who love coding and really dislike vibecoding, but atleast coding and prompting are the same mode of activity (typing in to a computer) so you don’t have to change your whole life to move from one to another.
I completely agree with you. The question is, are we headed to a situation where it is a new process that still allows the same expression, or is it fundamentally less able to facilitate that expression? we have to get beyond the one-shot style prompt shown here to something much more fine grained. The jury is still out for me on this, but I can believe we will get there. We just aren't anywhere close to it right now.
I’ll try through there but I’m more curious if there is a ChatGPT or Claude monthly subscription that is subsidized by investors to grow so I can play around more freely
The video quality is insane, but is there a model that doesn't make video where it looks like the characters are pausing at the end of their lines for a laugh track? There just seems to be an extra couple of beats after someone says something where they just stand dead still. What's the deal with that?
Video generation models generate videos of a predetermined duration, so if a character finishes a line but there's still seconds remaining, then the model still has to fill it in
it seems perfect for that - maintaining coherence over 30s-1min type window is achievable and so many ads are built on unrealistic premise to begin with.
To me, just reducing the cost alone seems secondary, far more important is putting the actual creative and marketing people directly in control of the output. Maybe the end result still gets sent to a pro studio for final production, but letting the true stakeholders directly create what they want could be a killer app. It may also be a terrible idea, like Homer Simpson's car - but that won't stop it being successful.
Yes, it is awesome. Funny we got this level of capability now and we all take it in our stride. Growing up in the 80's I loved futuristic scifi, feels like I am in it now. And even I am 'just' taking it in my stride.
Speaking of the eighties, special effects were still done without CGI and typically on the cheap. That didn't make the movies and television series any less fun to watch. In the end it's about what people do with the tools, not the quality of the tools. People commenting on how these generated videos are not perfect or are getting some details wrong, look at some blockbuster movies from last century. It's very easy to spot the special effects. And it did not matter.
But yes 80s science fiction is now science fact. Talking cars like in night rider are now definitely a thing. Some can even drive themselves. The robots in Buck Rogers, Star Wars, etc. look clumsy and fake compared to the real humanoid bots we are now starting to see. More close to I Robot, which when it came out 22 years ago was pure science fiction.
> That didn't make the movies and television series any less fun to watch.
Unfortunately it doesn’t apply to everybody :( Babylon 5 was mind blowing when I was a kid, and today the space and other sci fi parts from there, when I see them on YouTube, are quite cringe. Show is still great, but that’s the part where I’d be totally fine if some modern consistent ai re-rendered cgi to something more pleasant.
The quality of AI videos is blowing my mind. I can still see things that seem a bit "off" but it's hard to distinguish AI videos from real videos. Even blockbuster movies are starting to be dwarfed by what AI can create. How long until we see the first full length AI movie hit the theaters? No actors, no development team, just a guy prompting AI...
People said this kind of thing about MiniDV cameras, and then iPhones. The reality is that technology hasn't been a barrier for filmmakers for a very long time. Writing is still the hard part.
This will make longer (~30s) narrative add creation a lot better and more interesting at a reasonable price tag--roughly $7 best I can tell. Looking forward to trying it out.
Much like the widespread availability of cameras has caused a massive increase in photos of things that would otherwise not have been photographed, I think widespread availability of AI video/image generation will cause a massive increase in media from everyone who wants to show their thoughts to others. Of course the vast majority will be uninteresting slop, because most people just aren't very creative or original, but there will definitely be more chances for those who are. The popularity of AI parody videos is a sure sign of things to come.
If we pinned quality here and focused on cost and speed, video models could become fantastic general creative tools.
I think the emphasis on using them for ads & monetizable slop is because they’re currently too expensive and thus usually need to be part of a revenue generating pipeline.
Where can you actually get access to these models that isn't an outright scam? All the sites that promised to have Seedance 2 turned out to be scams. Does anyone know how to actually use it, and is it available to run yourself?
Is it just me or are these video models only actual use case is misinformation and spam? Sure they show us quirky and whimsy samples on the release page, but does anyone really believe that?
There are thriving Ai video communities who are trying to replicate big budget productions with indie resources. Look up Gossip Goblin. There's a parallel explosion of memes at the same time, like Balenciaga Harry Potter
So brainrot? I don't want to belittle anyone's creative pursuits, but really, I struggle to see how "Balenciaga Harry Potter" is any more artistic than "Italian Animals".
I think that will remain the use case of these tools as long as it is still so easy to tell they don't represent reality very well. Maybe that's a good thing, maybe we're not ready for a world where any video longer than a brief clip is indistinguishable from reality.
Also, much of the promotion of these tools come from a population of people unhappy with the current state of the real content creation industry that use real footage and real people because AI movie promoters see traditional media as mainly propaganda machines anyway, so these tools sort help level the playing field.
I belive thats what it WILL be, its heavily used in scammy and poltiical bullshit currently, but its only a matter of time until the cost vs benefit makes sense for hollywood and companies to start using it for actual commercials and tv shows and special effects etc for major videos.
Only if you don’t think about it for more than two or three seconds. Then you remember deepfake revenge porn and elon musk style child sexual material.
I get entertainment value out of watching the videos other haves created. I hear of educators creating educational value from the videos they make as it allows them to make higher quality explanations.
I've been having some similar thoughts lately about some application spaces being more harmful than others.
Coding, for example, seems somewhat benign, because code has to fulfil clear metrics: Either it works or it doesn't. Either it performs or it doesn't. As long as you are able to provide those metrics, you can to some extent treat it as a black box without losing much from a system view (lifecycle, long-term maintenance, keeping things working, broader qualified efficiency like re-use are of course other stories).
But on the other hand, delegating decision-making and thinking to models feels plain harmful to me. I have so many stories around me from office settings these days: "We have to pre-pone meeting XYZ, and we don't have enough time to prepare, so I made this AI analysis <power point deck>". This was then mis-prompted to fit a foregone conclusion, no one has time to do it or capacity to refute a 2000 word slop deck (except by slopping back), and crazy stuff becomes plan of record. This is really going to be a problem for organizations ...
That is an Indian English word, most non Indian English speakers are not familiar with it. reschedule or bring forward would be more appropriate for a global audience.
It's just you. Any director, storyteller, or filmmaker can see the potential here to realize scenes that simply wouldn't be possible otherwise. But indeed there's plenty of slop coming out too.
The quality is quite high, but an observation is the direction they're taking these models correlates heavily to the usage demand of China vs the West. Specifically, they're immensely focused on t2v for action / high effect shots. There's one human reference shot in the entire release page, and that one doesn't focus on dialog at all.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
It's not about cultural difference but simply what's easier.
Action sequences are forgiving. There's lots of motion and shiny vfx to fudge things over. Rapid cuts mean each shot can be short enough that the inevitable accumulation of AI hallucinations from frame to frame doesn't get distracting. Sets can be generic: if you're generating robots attacking New York, the viewer isn't going to have time to track if the bodega on the corner is in every shot, or if the Chrysler Building switches place around town. Shots can be practically from different cities and nobody will notice.
Carrying over an actor's performance is the opposite. You want long shots and impeccable scene stability. You can't just throw some additive-blended particles on the actor's face to distract from its generation deficiencies, like you can do for the Marvel scenes.
Bytedance is a corporation looking for clout, not to help filmmakers. They're releasing a model that makes them look good, and action sequences do that.
It's a counterintuitive thing to viewers who have long been told that vfx is the most expensive kind of moviemaking. With gen AI, these somewhat convincing Marvel pastiches are trivial to produce, but a sitcom episode is utterly impossible.
China has a population 4 times bigger than the US, and the films you cite did roughly 4 times better in each example.
But how much does ticket cost in China compared to the US? (Genuine question)
I hear this a lot, but I observed this is no longer that simple.
Like Green Book (2018) made over US$70.7 million in China.
And one of the most popular films in 2021 was Hi, Mom (2021), made over US$785 million.
Obsession (2026) is another one that is doing amazing in China, absolutely beating high action films like Supergirl ($12 million and going vs less than $1 million)
> Obsession (2026) is another one that is doing amazing in China
I haven’t seen the movie yet but I’ve been a subscriber to the YouTube channel of Curry and Cooper [1] for a long time. I’m looking forward to watching the movie.
I wonder how many of the other people that went to see Obsession were already fans of Curry and Cooper.
[1]: https://www.youtube.com/@thats_a_bad_idea
Supergirl is basically a lazy slob who is born with godlike powers, I can see why that doesn't play well in a culture that prioritises hard work and self-discipline.
Not doing terribly well anywhere as far as I've heard. That one is probably not cultural
Or western movies don’t have stories Chinese audience would relate to, but for dumb visual effect fests, there is no story so it works everywhere.
Are they targeting movies or ads?
hong kong films is the mirror image of hollywood and fantasy/action are a lot more popular than american.
But you're probably just blundering with value comparison since just per capita, your numbers arn't that' extreme, still the point stands.
I think it makes sense that historically American film focused on narrative based dramas and slow burn horror. Think of classic Hollywood hits such as Citizen Kane and Psycho. This is even apparent in the recent success of Obsession, which really follows strongly in the trend of Hitchcock, magical realism, human drama, psychological horror.
Compare this with historically successful Honk Kong films: Kung Fu Hustle, Police Story, Ip Man. I mean there are examples of Hong Kong films that are slow burn human dramas, like Chungking Express and Eat Drink Man Woman, and successful American action films like Diehard, but I think its clear that Hollywood film didn't start out with action movies, and Hong Kong film didn't start out with slow burn dramas, but adopted these genres afterwards; you could even claim that successful Hong Kong films like the Bruce Lee films of the 70s fueled the American appetite for action in the 80s, which was the most prominent era for that genre in the US.
As a corollary to this, it becomes clear that the Korean and Japanese film making industries are not as culturally distinct from Hollywood as Honk Kong film making is. For Korean film, the reason is obvious: Koreatown in Los Angeles is right next to Hollywood, in some respects is a part of Hollywood, and there are deep ties now between the Korean and American film making industries for that reason, and Korean-Americans have an overly high degree of representation in American media, and Korean productions often come to America to shoot (think of the end of Squid Games: it's almost certainly the case that they shot in LA because a producer has a cousin or something that works in Hollywood).
Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. But then I remember that I spend upwards of $10k on inference generating well over 50k images for storyboards, training models; and probably creating almost an hour of video (I assume). Yeah, I get that things can be economic if you don't use the latest models (ran some case studies on this), but the latest models are the most fun to work with. It doesn't scale as well as "vibe coding" stuff together on the weekend. And when things work really well its almost as if you're seeing an zoopraxiscope come to life for the first time; and you just want to keep going.
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
I went to Art Center for film. Loved it. But ended up writing software instead of shooting movies (while still also handling a lot of visual art direction, graphics work, UI, 3D animation, etc). Now I feel like we're starting to be roughly in the same boat as far as using prompts.
What bothers me is that every piece of content generated this way helps flood an already saturated market for content, while slowly degrading the expectations of what people see, to the point that no one will bother with shooting or animating anything anymore. Even if it's 50% worse, it's 90% cheaper, so the economics argur against producing any new physically made content. Simultaneously, it's cannibalizing all existing content. This points toward a feedback loop, like a snake eating its own tail. And even though Hollywood blockbusters have followed that pattern for a couple decades, it's demoralizing to me to see it enshrined as the future of film (or to hear from someone who makes films that it would be a preferred mode of creation).
IS that a real person's website? It looks entirely AI generated, even the copy and sample projects.
MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.
H3 results are underwhelming compared to LTX 2.3 so far:
https://www.reddit.com/r/StableDiffusion/comments/1vciy35/lt...
Perhaps refined ComfyUI workflows will squeeze more quality out of it, but it's definitely not in the realm of Seedance, and Lightricks is training LTX 2.5/"LTX-Next".
Note: Those results are a little misleading the good LTX results were generated with a good workflow in ComfyUI, and a prompt expanding local model.
Whilst the H3 result was generated with the raw api using his raw prompt.
If you ask ChatGPT or other AI model to "improve" your prompt (with cinematic, good lighting) generally, you will get also very good result from H3 also.
My own H3 test show that H3 (API edition) is undoubtedly better than LTX in prompt adherence ! The real comparison will be with the edition of H3 we get to run locally.
Awesome! Have been a bit in the dark of the latest models, will have a look. I am due an upgrade for my local machine and GPU, so that excites me as well.
I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models.
It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)
Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.
Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.
Thanks for the input! And no I agree with your assessment of the website, I have some plans of a much simpler redesign soon, and I definitely value that critique. I think my "digital garden" has been through at least a few dozen of iterations, and sometimes I just get too carried away with it. The good part being, the next iteration always starts out better than the previous one (or at least I hope so :)).
There is a woman on twitter who makes seedance videos of her and Dario from anthropic, they're kinda weird but generally pretty high quality: https://x.com/CuiMao/status/2058458683781365873 (full collection: https://x.com/CuiMao/status/2082740754380984373) - seeing them was the first time I'd been impressed with AI video gen.
For first of us without twitter:
https://xcancel.com/CuiMao/status/2058458683781365873
I can not believe how good the AI video has gotten. In one spot the written letters were obvious AI. Other than that - I couldn't find a thing.
the thing that upped the immersion was the the way it emulated the optical zoom when switching from the 1x lense to the 3x lense
There were tons of continuity errors, but I suppose bad film makers make those too and often get away with it. The worst was the gift that was boxed in one shot, then open and containing a gift that didn't fit in the box. Another more subtle one was after they bumped into each other they were standing in front of a wall suggesting she just ran out from the wall.
This is deeply disturbing…for so…so so many reasons.
This is absolutely hilarious
https://x.com/CuiMao/status/2049828401201246395
This specific video is the first time I have been distributed by AI video.
This is extraordinary lifelike.
There's honestly a good chance you have seen an AI video in the last several months without realizing it. AI images and videos got crazy good suddenly
disturbed?
The future is here. It’s just not evenly disturbed.
distributed!
It seems like it was clipped from this wild 4min video's ending: https://x.com/CuiMao/status/2058458683781365873
Super strange. It feels like something a stalker would do...
I dont normally like AI videos but this was just amazing, by next year the tech will be perfect, cheaper, we are going to AI videos everywhere whether anyone likes it or not.
people like it. Maybe not older generations, but we're about to get a fresh new generation of people that will not have a lifetime of experience looking at human generated content, and so have no real bias against AI content.
And then every generation after that will be born into increasingly a AI-video rich world.
The older generation loves them the most. Sometimes my mom sends me several AI generated reels in one day and she doesn't care if it's AI or not she just finds them funny.
Going to be a wild future that is for sure. I can see many ways this will go, some not so good.
same here.
Also in my country there is trend of remaking the old songs, with AI.
modernised sound with AI voice singing and AI video. They have ton of views in YT.
This one posted above looks fine to me, even after five decides of looking at human generated content:
https://xcancel.com/CuiMao/status/2058458683781365873
Honestly, it is frightening.
> And then every generation after that will be born into increasingly a AI-video rich world.
So 2-3 generations if we're lucky.
5 years. Most people do not care where what they consume originates (out of sight, out of mind) as long as it's highly personalized and abundant
Older generations absolutely love AI slop. If you haven't used Facebook recently, it's all old people sharing fake videos of animals doing silly things.
Well I find that after a certain age the behavior of old generations reverts to the same childlike patterns of young generations. It’s always the people in the middle who are getting screwed, they have to be the adults.
Right now we're going through multiple concurrent but slightly non-overlapping transitions where AI is almost good enough for something, and seeing a predictable surge of spam/bad quality crap people don't like, followed by a whiplash effect when it reaches human or superhuman levels of performance.
I think most developers would agree that the Cursor vibe-code era sentiment towards AI for coding was right in the sense that it wasn't that useful then, and a lot of people did and do stupid things with it, but it really did deliver quite substantially on its promise just a couple years after the "consensus" was often "that's never going to work".
It's very strange to me how consistently public opinion shifts on this problem. In the long run we're all dead but whether it takes one or five years, you probably will be there to live through it. I assume whoever is downvoting you is just thinking with their consumer-brain or "I don't want that" (the way it is now) rather than that this will be no different than video games, or wasn't a child who watched youtubepoop or other weird stuff because it was funny.
Nah.
Sad. Wish there was a huge library of human generated content they can look at.
I wonder if it'd handle chopsticks better: https://www.dropbox.com/scl/fi/p0yrn0x4q5pdz1ibl0gfa/Screens... (from your first link)
People are still figuring out the social norms surrounding "deepfakes" and I'm convinced there are some versions of the future where it becomes normalized for content under essentially fair use/free speech doctrine, and many where it's used to restrict free speech in the small number of jurisdictions where public figures would currently be allowed to be featured in this content.
Obviously there are a lot of ways it shouldn't be used but I want to live in the future where it's something we use to have fun, where reasonable, rather than a dangerous/sketchy taboo
That is absolutely hilarious. Dario Amodei, a romcom lead.
The whole damm Makoto Shinkai based dorama style too!
This is blessed and cursed timeline at the same time.
There is a ton of vertical dramas now (Duanju) made purely of AI. Even on chinese streaming services you find it.
I can't make up my mind if the simplistic narrative typical of Chinese romcom shorts is more funny or more weird.
The scene that says “Stanford University” looks a lot like UCLA.
UCLA is often used as a stand in for many colleges in movie sets (notably Harvard, as one of the buildings in UCLA has the Harvard crest on it specifically for this purpose) due to its proximity to Hollywood. It likely ended up in the training data that way
Yep, I notice it all the time in movies and tv shows. They were filming the movie "Old School" my freshman year there, I thought it was so cool. Makes sense it is in a lot of the training data
I guess the orange sphincter is really a flower…
"Intriguing... but highly disturbing"
Not a fan of this AI slop. We should boycott it
That battle is already lost
Seedance 2.5 looks amazing, but MiniMax H3 is going to be open weights within 24 hours: https://fal.ai/minimax-h3. According to the ComfyUI team, it should even work acceptably on mid-range consumer GPUs like the 3080.
I'd honestly take the slight quality hit for more control and lower costs.
This. I'm very much happy to see minimax-h3 be released. I have so many projects that I have on the back burner that I could complete within minutes instead of hours and hours. "Typography, UI, and Graphics".
New model day, here we go again! It’s a week of madness and all the wrong advice… and at the end of the week, the model is either dead or a massive hit. There is little middle ground.
I’m watching H3 to see if it fixes the nonsense that LTX introduced and what happens next… WAN is due, LTX says a new model is coming, Flux3 does video.
It went from zero video models to a lot of options due this year.
In second video, Look at the reflection of the staff member headset :)
Such models will have fascinating use-cases. But "generating realistic content" is a comedic resemblance of realism movement in painting.
I don't think audio, image or video generation should exist. I haven't seen enough positive applications to justify the amount of harm these tools are being used to cause.
What are the positive applications? This all seems to be generating fake images/videos just because we can, not because it solves any real problems. It makes me wonder where the money is coming from to fund this stuff. Why would a regular person want to generate fake videos except for purposes of deception or debauchery?
The lack of creative understanding is astounding: all these generative media AI are fantastic for education applications. Sure, the vast majority want to use them to jerk off, but that's what humanity does anyway. Education is where these are extremely useful, and that is for the aid in conveying understanding to those that need the Jazz Hands of an image or audio or a video presentation to grasp the concept they are not.
Benefit: entertainment
Downsides: *
(Well <all> minus entertainment)
While I generally think it's balanced against a lot of negative applications (misinformation, deepfakes, etc.), there are some good applications. I think regulation just hasn't sufficiently caught up with the harm they can cause.
To give a good use case (audio generation specifically), the 2002 game Morrowind has had a resurgence lately and it's a really great game, but it lacks fully voiced lines for dialogue. A community project called Voices of Vvardenfell uses ElevenAI to generate voice lines from the voiced lines that do exist in the game. They don't take away jobs from artists because it was a monumental (unpaid) undertaking from community volunteers to do this before (there was such a big project but it was abandoned last I saw).
So as with any emerging technologies, I think it just needs to be properly regulated.
> They don't take away jobs from artists
Does the game with improvements now better compete against other games with paid human voice actors?
The game is from 2002. Anybody wanting to play the game either buys it for nostalgia or as an enthusiast. All the mod does is increase accessibility.
Morrowind wasn't meant to be narrated, NPCs blur out walls of text with hyperlinks in them leading to more walls of text. If I had to sit idle while someone/something was reading all of this to me, I'd be bored out of my mind.
What "harm"?
The only "harm" is from the authoritarians wanting to ensure only they have the means to produce propaganda for the masses.
Now the masses can amplify their own voices instead.
Some voices are “$politicalParty should win at all costs” and “my ex-spouse should be humiliated” and “my boss should go to jail for the pittance they pay me”, and instead of text opinions posted or voices shouted, there’s video gen to amplify them. They didn’t have the skill to make a convincing fake video themselves but if tools are at their fingertips… temptation.
Really not ideal although obviously amazing artists will do incredible things with the technology that many of us will enjoy.
Cat out of bag, of course.
This is what the First Amendment protects. True freedom.
It's going to be just like social media. A flood of content from the talentless masses exposing us to their lazy and shallow ideas they just cannot keep to themselves. Once in a while a diamond in the rough will be found, but at what cost?
The long tail dream didn’t really happen in any single market which was democratized the same way in the past. Why do you expect something else now, especially on an already oversaturated market?
How does creating fake video amplify anyone's voice?
Now they cause harm because there's people that may believe a fake video is real. But people is developing skepticism about what they see in videos and learning to think that they may be generated, before taking them as true.
What kind of harms do you have in mind?
No longer being able to tell whether any image or video is real due to the extreme ease of creating realistic fakes, is itself a harm.
It will be solved. Walleted, vertically integrated Apple can do a digital authenticity signature and guarantee that the video you see was at least shot on iPhone, and I won’t be surprised if it will happen soon. This + some lidar info will be enabled and everybody except minority will say that this is a great thing to have.
until models will be trained to fake that too
You've seen movies before. You know that they can make fake videos look realistic. We have an entire industry built around it, so why is it only a harm now?
Since it took a certain amount of time and effort, as well as skill in the past, it was a significantly lesser problem than it's becoming now. Especially with video footage.
plus all the money and people involved created a paper trail, which meant proving something was "faked" much easier
that no longer exists with genAI models
It might actually be helpful to push through and out to the other side.
Think about how many completely fake videos are going to be flying around for the next election. Nobody is going to believe anything the see online.
What is the the next stage, where the only things you believe are things you see with your own eyes.
>things you see with your own eyes
Oh I wouldn't trust myself!
https://en.wikipedia.org/wiki/Eyewitness_testimony#Reliabili...Also-mentioned concepts: misinformation effect, source misattribution, confirmation bias.
> I haven't seen enough positive applications
If in the future they can create a reward model to aid with creating super human levels of appeal tailored to me, I would considered it one of the best technologies ever made. I'm not sentimental about the source, pretty videos and pictures make me happy
Fyi based on what people are saying on twitter, seedance 2.5 is ~2x as expensive. For a 30 second generation at dreamina it costs 1440 credits or around $15
Which might mean a lot for individual but is peanuts for a business
I was trying to get at that its 2x more expensive for businesses, and personally I dont think 2.5 is worth double the variable cost. At least initially
Which is next to nothing compared to the cost of a movie, animated TV show, or quality YouTube video.
Wait, but you compare the cost of animated show or a movie which is in its final, acceptable quality!
The cost of that $15 per 30sec is NOT FINAL - you'll pretty much end up deciding you want another take - either due to some inconsistencies, non-physical movements, or anything else, which screams "that's fake".
So the cost is most probably much higher; how much higher? We don't really know, as there are virtually no acceptable shows created by that tool yet (until there is something I don't know about).
>> The cost of that $15 per 30sec is NOT FINAL - you'll pretty much end up deciding you want another take - either due to some inconsistencies, non-physical movements, or anything else, which screams "that's fake".
Thats true for grok and other models. You need to generate endlessly and cut endlessly and stitch them up together.
but for seedance you pretty much get stellar output which is extremely consistent in a single shot. Like there is a very good chance you will like what you see.
My experience with AI image generation is---
If you want something vague and generic and don't particularly care about the details, then maybe you can get want you want on the first try, but you'll likely just get what everyone would call "slop".
If you have something very specific in mind, then expect to spend several dozen generations before you get what you want, or something close to it, and then you may still have to spend more time making final adjustments manually in an image editor.
...and this is just for one image. I have not tried AI video production yet but I imagine it would involve considerably more effort.
$15/30s is $1800/hr, so assuming 1/10 generations are good, that's $18000/hr for final content. A few orders of magnitude below professional movies for sure, but still not trivial.
It is not a movie though. It's a take.
The video is silly but the quality is very impressive. We are very close to climbing out the other side of the uncanny valley.
I would say we are pretty much there.
"Pretty much" is what the valley means though :)
It has improved a lot, but these demo reels still have all AI video issues. Flash cut salad (including the scenes that should have longer cuts), unnatural motion that looks animated, unprompted YouTube-face acting, etc. Admittedly it's all a lot less pronounced in this version.
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
This seems extraordinarily good to me, compared to what I've seen before. Their washing machine advert example seems like it's as good as anything else on social media. I'm shocked by the quality and the coherence they're able to maintain, I assume it's really good at using those reference images they mention in their prompts. I think the only one that's noticeably bad is the concert hall, where the first few seconds show almost empty stalls and then toward the end it shows a full audience, and that audience also looks a bit off.
The sample video looks incredible and the ability to keep details consistent for so long is impressive but it _still_ looks screams of AI in every single shot. I can’t put my finger on why, something about the way the rooms are put together, the expressions of the faces, the movements… It’s very unsettling.
I think we as humans are just really good at being able to tell. CGI has been around for decades and yet we can still look at it and say "yeah that was CGI". Uncanny valley I guess?
No need to fret, this is the worst it's going to be. Remember, Will Smith + Spaghetti videos?
The bicycle was improved a lot when it was created but did changed a little since then too
yeah we got automobile, motorcycles, planes, lot of stuff since then
It could also still be the best it's ever going to be.
Anybody old enough to remember when https://imagen.research.google/video/ was SOTA?
do you remember that article from google where it would hallucinate as if it was on acid/psychedelics and people going crazy saying this is the AI "imagining reality" and then shortly after some Googler went on a very brief media tour saying LLMs were sentient?
its crazy how much people extrapolate, the consensus back then was "cute but we won't get there" then suddenly we got GPT 3.0, 4o, 5.x, seedance
what people really underestimate is how much faster progress is now with AI
That phase was amazing, lots of crazy animal faces in iridescent paisley :)
Unfortunately trillions of dollars have been invested to ensure that's not the case..
unlikely unless they ban capitalism
You know what, in my previous life as a filmmaker I could've only dreamed of such a thing. Filmmaking is an art form which you cannot do alone. Outside of a lot of time and money, you need cooperation of a number of people and each day of production you end up accumulating a set of compromises to your vision. Your taste is what makes you tolerate that or not, and it's exhausting. It matters if you're the author since at the end of the day it's your name on it, not the crew (as much).
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
Weird-ass times.
I think for me it's just missing the "essence". Making a film is hard, but that labor with multiple people who are passionate about the project is just as important as the end result.
I've been trying to generate concept art for a game idea. I figure, I can't draw so why not have something else bring my vision to life. After dozens of failed attempts and burned tokens I decided to just hire someone. The experience was night and day. They provided thoughtful details to help tell a story in the image. They got to tell me what they added and why, etc. This isn't even their project but their passion for it showed in the end result.
These models are impressive, but I think people watch movies because of the essence as much as the movie itself.
This is the same level of arguing that only painting is the real art and photography is just clicking the button
Was just going to say... Even if all AI gives me is the opportunity to bring my visions to life even somewhat terribly... That's a lot better than I'd be able to do without it. Which is for them to go nowhere. I'll take it.
I am not a film maker but I had some sketches on my head that I always wanted to try to make. Even Gemini scratched the itch for me, and my folder “projects” is growing with other ideas that I’d like to sketch
The thing is, it has the same essential property that you don't control it any more. The problem with these tools is the lack of fine grained control means there is no room for you to express your individual creative input. Your total input is a few sentences of a prompt, then the AI did all the creative part. If the creative input is what you enjoyed, it is actually not that much more (or even less) here than it was with traditional film.
It will change soon. Ai is enabler. Again, as with photography and especially digital photography- everybody can click button but not everybody is photographer. Ai helps to raise the floor, not the ceiling
This is a matter of effort and direction and not an inherent flaw of the tools. Look at Apple's recent image tools to change perspective. Tools to improve artistry are improving.
Could you tell more? Is it new iOS features?
Another aspect is that if you struggle and struggle and create something great, people can see the work you put in, and you'll visibly stand above the rest who don't have the same level of perseverance. But if anyone can create big studio quality visuals with a prompt, how do you stand out? "When everyone's super, no-one will be."
I often see this kind of passive-aggressive characterization of the process with AI that's meant to bait people who obviously are expressing their creativity with these models.
I hope no one will fall for it: even when you emphatically reply "I spent the last year building the pipelines I use" it'll still be treated as equivalent a couple of sentences relative to the old ways.
-
At the end of the day, the painful lesson some people are (re)-learning is that creative expression doesn't have to be tied to any specific process. A lot of people fall in love with creative expression + process... but some fall in love with just the process, and some fall in love with creative regardless of the process.
AI does not favor those of us who loved the process (that's me with coding), but it's definitely capable of allowing new people to experience what it's like to have creative input in your mind and see it expressed outside of your mind.
Often in making art you need to love the process. Otherwise you’ll never get good enough. Most filmmakers don’t even watch their films after they have spent years with it discussing, planning, shooting, editing etc. There is lots of travelling, new friendships, comradery involved and in a way it becomes the whole scaffolding of your life. Replacing this with prompting is probably not interesting to this particular group of people, but might be appealing to completely new set of people.
There are people who love coding and really dislike vibecoding, but atleast coding and prompting are the same mode of activity (typing in to a computer) so you don’t have to change your whole life to move from one to another.
I completely agree with you. The question is, are we headed to a situation where it is a new process that still allows the same expression, or is it fundamentally less able to facilitate that expression? we have to get beyond the one-shot style prompt shown here to something much more fine grained. The jury is still out for me on this, but I can believe we will get there. We just aren't anywhere close to it right now.
I am going through the same struggle but with software development.
Where can you use seance to make movies? Do you have to sign up to a. Third party? Do they have an official app I don’t know about or is there an api.
I just can’t figure it out!!
There are no stupid questions only stupid people.. when it comes to figuring this out I’m one of those people
AFAIK there are providers that give you API access to Seedance, like OpenRouter etc.
I’ll try through there but I’m more curious if there is a ChatGPT or Claude monthly subscription that is subsidized by investors to grow so I can play around more freely
From the article:
> Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk.
Jimeng is https://jimeng.jianying.com/ (I'm pretty sure, there are a lot of copycats, but this one uses Douyin SSO for login, so seems legit) and available in English as "Dreamina" https://dreamina.capcut.com / https://play.google.com/store/apps/details?id=com.lemon.drea... (note that it's made by Bytedance; on iOS the search results seem to be copycats only) Doubao is https://www.doubao.com / Dola https://www.dola.com in English. There seems to be no way to find out how much a subscription costs without making an account.
The video quality is insane, but is there a model that doesn't make video where it looks like the characters are pausing at the end of their lines for a laugh track? There just seems to be an extra couple of beats after someone says something where they just stand dead still. What's the deal with that?
Video generation models generate videos of a predetermined duration, so if a character finishes a line but there's still seconds remaining, then the model still has to fill it in
Have already seen countless major ad spots using ai generated videos; big name companies. One major use case.
it seems perfect for that - maintaining coherence over 30s-1min type window is achievable and so many ads are built on unrealistic premise to begin with.
To me, just reducing the cost alone seems secondary, far more important is putting the actual creative and marketing people directly in control of the output. Maybe the end result still gets sent to a pro studio for final production, but letting the true stakeholders directly create what they want could be a killer app. It may also be a terrible idea, like Homer Simpson's car - but that won't stop it being successful.
What language is the girl singing in?
Yes, it is awesome. Funny we got this level of capability now and we all take it in our stride. Growing up in the 80's I loved futuristic scifi, feels like I am in it now. And even I am 'just' taking it in my stride.
Speaking of the eighties, special effects were still done without CGI and typically on the cheap. That didn't make the movies and television series any less fun to watch. In the end it's about what people do with the tools, not the quality of the tools. People commenting on how these generated videos are not perfect or are getting some details wrong, look at some blockbuster movies from last century. It's very easy to spot the special effects. And it did not matter.
But yes 80s science fiction is now science fact. Talking cars like in night rider are now definitely a thing. Some can even drive themselves. The robots in Buck Rogers, Star Wars, etc. look clumsy and fake compared to the real humanoid bots we are now starting to see. More close to I Robot, which when it came out 22 years ago was pure science fiction.
> That didn't make the movies and television series any less fun to watch.
Unfortunately it doesn’t apply to everybody :( Babylon 5 was mind blowing when I was a kid, and today the space and other sci fi parts from there, when I see them on YouTube, are quite cringe. Show is still great, but that’s the part where I’d be totally fine if some modern consistent ai re-rendered cgi to something more pleasant.
The emotional delivery is so flat and the words are phrased so awkwardly. They're still a lot of work to be done in this space before its believable.
Where does the sound on these AI videos come from, is that AI generated in the same pass, a second model, or is it added by a human after the fact?
Been using this and it’s really a phenomenal model, big step up in quality and capability.
This feels like the serious inflection point for high quality full length feature film productions using this tech.
The quality of AI videos is blowing my mind. I can still see things that seem a bit "off" but it's hard to distinguish AI videos from real videos. Even blockbuster movies are starting to be dwarfed by what AI can create. How long until we see the first full length AI movie hit the theaters? No actors, no development team, just a guy prompting AI...
People said this kind of thing about MiniDV cameras, and then iPhones. The reality is that technology hasn't been a barrier for filmmakers for a very long time. Writing is still the hard part.
I would like to see seedance 2.5 at 1080 vs seedance 2.0 but at 4k to see which models outperform
From our experience putting in more money leads to more computerand better performance from the prompt
The generated content is close to real.
This will make longer (~30s) narrative add creation a lot better and more interesting at a reasonable price tag--roughly $7 best I can tell. Looking forward to trying it out.
Who here is using Jimeng AI or Doubao Pro? Is it comparable to Claude / Gemini / GPT?
https://news.ycombinator.com/item?id=49122670
Much like the widespread availability of cameras has caused a massive increase in photos of things that would otherwise not have been photographed, I think widespread availability of AI video/image generation will cause a massive increase in media from everyone who wants to show their thoughts to others. Of course the vast majority will be uninteresting slop, because most people just aren't very creative or original, but there will definitely be more chances for those who are. The popularity of AI parody videos is a sure sign of things to come.
Weird there is no option to sign up and try the model out.
Its not released in the US yet
If only I was rich...
interesting to see more progress
This looks insane!
If we pinned quality here and focused on cost and speed, video models could become fantastic general creative tools.
I think the emphasis on using them for ads & monetizable slop is because they’re currently too expensive and thus usually need to be part of a revenue generating pipeline.
We are officially one step closer to Interdimensional Cable!
Where can you actually get access to these models that isn't an outright scam? All the sites that promised to have Seedance 2 turned out to be scams. Does anyone know how to actually use it, and is it available to run yourself?
Dreamina (https://dreamina.capcut.com/) is the “official” access for this.
IIRC they give Seedance 2.5 yesterday. Byteplus API is the official way to access it as an API but I am not sure if they serve 2.5 yet.
They have not. It also seems the dreamina acess is delayed for the US
Openrouter has video generation also. They do have Seedance which afaik routes via something called ‚atlas cloud’, but it did work.
Yeah wondering that. I ended up almost paying for one when 2.0 came out. Still not sure I've found something official.
23 Jun 2026 03:31:03 UTC
Seedance 2.5, generating a complete 30-second video in one go
https://twitter.com/xiaohu/status/2069236441964818508
https://news.ycombinator.com/item?id=48639900
[ok]
23 Jun 2026 03:53:32 UTC
Seedance 2.5 pushes AI video toward references and control
https://news.ycombinator.com/item?id=48640057
[flagged] [dead]
23 Jun 2026 04:47:30 UTC
SeeDance 2.5 Is Stunning
https://twitter.com/Long4AI/status/2069262125776920582
https://news.ycombinator.com/item?id=48640441
[ok]
23 Jun 2026 13:50:27 UTC
ByteDance's Seedance 2.5 breaks the 30-second barrier for AI video generation
https://the-decoder.com/bytedances-seedance-2-5-breaks-the-3...
https://news.ycombinator.com/item?id=48644976
[ok]
23 Jun 2026 15:48:31 UTC
Seedance 2.5
https://seedance2.ai/seedance-2-5
https://news.ycombinator.com/item?id=48646930
[ok]
24 Jun 2026 08:01:56 UTC
Seedance 2.5
https://www.seedance2ai.app/tools/seedance-2-5
https://news.ycombinator.com/item?id=48656671
[ok]
25 Jun 2026 01:29:28 UTC
Show HN: Seedance 2.5, ByteDance's upgraded multimodal AI video generator
https://seedance25ai.cc/
https://news.ycombinator.com/item?id=48667707
[flagged] [dead]
29 Jun 2026 09:32:30 UTC
Seedance 2.5
https://news.ycombinator.com/item?id=48716942
[flagged] [dead]
31 Jul 2026 13:09:53 UTC
Seedance 2.5
https://seed.bytedance.com/en/seedance2_5
https://news.ycombinator.com/item?id=49122670
[ok]
Why amplify the copycat spam domains?
The first video is wild. I thought the office workers were real. Jesus Christ.
I'm looking forwarding to AI-based video refactoring for when you want to tweak them.
amazing
Jesus christ, did you people watch those videos! there so long and coherent and WOW.
They’re*
Sorry
Is it just me or are these video models only actual use case is misinformation and spam? Sure they show us quirky and whimsy samples on the release page, but does anyone really believe that?
There are thriving Ai video communities who are trying to replicate big budget productions with indie resources. Look up Gossip Goblin. There's a parallel explosion of memes at the same time, like Balenciaga Harry Potter
So brainrot? I don't want to belittle anyone's creative pursuits, but really, I struggle to see how "Balenciaga Harry Potter" is any more artistic than "Italian Animals".
I think that will remain the use case of these tools as long as it is still so easy to tell they don't represent reality very well. Maybe that's a good thing, maybe we're not ready for a world where any video longer than a brief clip is indistinguishable from reality.
Also, much of the promotion of these tools come from a population of people unhappy with the current state of the real content creation industry that use real footage and real people because AI movie promoters see traditional media as mainly propaganda machines anyway, so these tools sort help level the playing field.
I belive thats what it WILL be, its heavily used in scammy and poltiical bullshit currently, but its only a matter of time until the cost vs benefit makes sense for hollywood and companies to start using it for actual commercials and tv shows and special effects etc for major videos.
Of course not. You can use them to make objectively silly fetish fodder, too.
P*rn!
Unironically, it might be more ethical.
Only if you don’t think about it for more than two or three seconds. Then you remember deepfake revenge porn and elon musk style child sexual material.
What about "beefy animal men pin-ups" and "ideal body self-deepfakes"?
I get entertainment value out of watching the videos other haves created. I hear of educators creating educational value from the videos they make as it allows them to make higher quality explanations.
Unlike the coding models it seems like the generative art models do a lot more harm than good.
I've been having some similar thoughts lately about some application spaces being more harmful than others.
Coding, for example, seems somewhat benign, because code has to fulfil clear metrics: Either it works or it doesn't. Either it performs or it doesn't. As long as you are able to provide those metrics, you can to some extent treat it as a black box without losing much from a system view (lifecycle, long-term maintenance, keeping things working, broader qualified efficiency like re-use are of course other stories).
But on the other hand, delegating decision-making and thinking to models feels plain harmful to me. I have so many stories around me from office settings these days: "We have to pre-pone meeting XYZ, and we don't have enough time to prepare, so I made this AI analysis <power point deck>". This was then mis-prompted to fit a foregone conclusion, no one has time to do it or capacity to refute a 2000 word slop deck (except by slopping back), and crazy stuff becomes plan of record. This is really going to be a problem for organizations ...
> pre-pone
That is an Indian English word, most non Indian English speakers are not familiar with it. reschedule or bring forward would be more appropriate for a global audience.
I'm not Indian, and I see it used quite a bit by non-Indians :-)
I also don't think it's difficult to understand given the prevalence of "post-pone".
It's just you. Any director, storyteller, or filmmaker can see the potential here to realize scenes that simply wouldn't be possible otherwise. But indeed there's plenty of slop coming out too.
No weights, no care.
Perfect for generating the mindless garbage that tiktok junkies are dependent upon.