I recall how there were custom chips made for LLaMA for that make it super fast which is cool but am I wrong for feeling like that would slow down the release schedule of new frontier models? Couldn't tell from the article if these would be ASICs or more general purpose.
ASICs may be too slow to iterate. If we want to close the development loop faster, I think FPGAs are a way to do so. With everything moving so fast nowadays, by the time the ASICs are ready for production, they could be already outdated, so FPGAs could be a quick way to test some novel designs, with or without LLMs' help.
If the ASIC runs a model 100x better (faster, cheaper, less energy, etc.), it will remain competitive on any task for which frontier models don't improve as quickly. Any task for which current models are "good enough" arguably fall in this category. We just need to cross a threshold in terms of how many such tasks there are and their economic value. Also worth pointing out is that there's probably a lot of latent uses of LLMs, currently unexplored, that would be enabled specifically by a very low latency. Right now, I'd feel comfortable taking the risk.
Good, can they please make their own RAM as well and leave some for the little people?
I recall how there were custom chips made for LLaMA for that make it super fast which is cool but am I wrong for feeling like that would slow down the release schedule of new frontier models? Couldn't tell from the article if these would be ASICs or more general purpose.
ASICs may be too slow to iterate. If we want to close the development loop faster, I think FPGAs are a way to do so. With everything moving so fast nowadays, by the time the ASICs are ready for production, they could be already outdated, so FPGAs could be a quick way to test some novel designs, with or without LLMs' help.
If the ASIC runs a model 100x better (faster, cheaper, less energy, etc.), it will remain competitive on any task for which frontier models don't improve as quickly. Any task for which current models are "good enough" arguably fall in this category. We just need to cross a threshold in terms of how many such tasks there are and their economic value. Also worth pointing out is that there's probably a lot of latent uses of LLMs, currently unexplored, that would be enabled specifically by a very low latency. Right now, I'd feel comfortable taking the risk.
The tech that makes sense here is structured ASIC
Qwen 3.8 came up this week with claims on being good for designing chips. so competition has to tighten.
Good, another way they can waste more money & hopefully crash faster & harder.
Now they can claim they're "full stack" like Google as another justification for their $1 trillion IPO.
duh. they all are.
They should vibe code it and see how it goes.
Translation: Anthropic is taking on more debt.