I just bought a used Ayaneo 2 handheld, it has an old(er) mobile RDNA 2 GPU and I was blown away by how well this thing performed under Linux. Almost everything (that's not a recent AAA game) runs beautiful and a lot faster/smoother than it does under Windows. The experience has been so good that I'm considering switching my main pc (with a 9070XT) to Linux as well.
No doubt Timur contributed heavily to this given Valves Steamdeck (which uses a very similar but slower GPU).
Given the current hardware prices it's pretty awesome to see someone squeezing maximum performance out of old hardware!
I'm sure Valve probably still deserves the credit, since the Steam Deck uses a RDNA 2 gpu, but the article mentions that Timur specifically worked on gcn1.0 GPUs - e.g. radeon HD 7800/7900 from 2012 (!)
My experience has been that for Linux you are always better buying older mid tier hardware because all of the issues and optimisations have already been worked out and those changes have flowed through to your distro so you aren’t waiting on kernel updates.
I've switched my gaming PC last year to Linux and it's been flawless.
Originally ran a 3070 which matches your "older mid tier" description and then upgraded to a 9070XT maybe 6 months after launch so not so old or mid-tier.
My gaming PC has a 9070 non-XT, runs Linux, and runs all but a small handful of games [0] just fine. I suspect you'll be happy with the switch.
I run games both new and old and one of my favorite guilty pleasures is looking at the Steam forum for a four, eight, or ten year old game that I've started playing because it has recently become popular again and reading the complaints from Windows users about how a driver update screwed up the game... whether because of glitchy or incorrect graphics or unavoidable crashes. Meanwhile, I'm cruising along on Proton with zero issues. :smug-face:
[0] ...that small handful includes those that go out of their way to be incompatible with Proton...
Use as a dedicated GPU for encoding and decoding video.
Post processing like frame interpolation or superresolution. Use for GPGPU workloads. Run additional monitors independently. Use for GPU passthrough to virtual machines. Use as a backup GPU for troubleshooting. Use for test code without breaking the main GPU
Even proprietary drivers can use Mesa. Most of Mesa is MIT-licensed.
It's just that Nvidia doesn't care much about Linux, and -as always- Nvidia ignores what everyone else is doing and does their own thing. Sometimes doing their own thing works out really well in the short run, but -long term- they always fall behind.
One thing they should've learned from Nvidia is that it's really worth it for them to invest making their devices function as broadly as possible. Crypto and AI waves both benefited Nvidia much more than AMD partly due to their devices being more universally usable.
That "partly" was worth hundreds of billions of dollars but AMD were cheap/shortsighted enough to hire a few dedicated engineers.
Perhaps René, the maintainer of T/2 Linux, will end up vibe coding an NVIDIA driver one of these days. He’s been reverse engineering and vibe coding a bunch of drivers for old graphics lately and live streaming everything.
Maybe there’s a prestige angle to somewhat shame them into keeping up with nvidia. Though my sense is that they’re barely peers anymore in this space. :(
Honestly, the folks writing inference drivers for Llama.cpp / GGML would sort of benefit from better compiler work like this.
In general, Valve's work has been exceptional and supplementing AMD's own ROCM/OpenCL and Vulkan team, they've gotten a lot of defaults right and people should work together with Valve to improve support for their chips.
Most of these older cards are lacking the physical hardware for fp8 or other lower precisions that most quantized models use. Or the memory to run models at higher precision.
I have a Radeon RX 6900 XT (16 GB VRAM, originally released in 2021), and it's possible to run some lightweight models to have okayish performance and quality of output, but nothing I've tried has come anywhere close to the quality even of the models I can use for free from OpenCode Zen or the free tier of Openrouter. If you want to keep everything local on the same card I have, it requires putting up with a model that's noticeably worse in virtually every metric than what you can get for free elsewhere, and the GPUs this article are talking about are three times as old as mine.
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
I have no clue, and I agree that it does not seem sustainable. Either someone needs to find a magic solution to making it a lot cheaper, or a lot of companies are going to need a lot of money from somewhere that isn't clear.
With an extra 8gb of vram you could run qwen 3.8 27b pretty comfortably, which isn't quite as good as frontier models but definitely on par with free models on openrouter and whatnot.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
I just bought a used Ayaneo 2 handheld, it has an old(er) mobile RDNA 2 GPU and I was blown away by how well this thing performed under Linux. Almost everything (that's not a recent AAA game) runs beautiful and a lot faster/smoother than it does under Windows. The experience has been so good that I'm considering switching my main pc (with a 9070XT) to Linux as well.
No doubt Timur contributed heavily to this given Valves Steamdeck (which uses a very similar but slower GPU).
Given the current hardware prices it's pretty awesome to see someone squeezing maximum performance out of old hardware!
I'm sure Valve probably still deserves the credit, since the Steam Deck uses a RDNA 2 gpu, but the article mentions that Timur specifically worked on gcn1.0 GPUs - e.g. radeon HD 7800/7900 from 2012 (!)
My experience has been that for Linux you are always better buying older mid tier hardware because all of the issues and optimisations have already been worked out and those changes have flowed through to your distro so you aren’t waiting on kernel updates.
I've switched my gaming PC last year to Linux and it's been flawless.
Originally ran a 3070 which matches your "older mid tier" description and then upgraded to a 9070XT maybe 6 months after launch so not so old or mid-tier.
Steam Deck amazes me, I've had it for years and it doesn't feel dated in the slightest.
My gaming PC has a 9070 non-XT, runs Linux, and runs all but a small handful of games [0] just fine. I suspect you'll be happy with the switch.
I run games both new and old and one of my favorite guilty pleasures is looking at the Steam forum for a four, eight, or ten year old game that I've started playing because it has recently become popular again and reading the complaints from Windows users about how a driver update screwed up the game... whether because of glitchy or incorrect graphics or unavoidable crashes. Meanwhile, I'm cruising along on Proton with zero issues. :smug-face:
[0] ...that small handful includes those that go out of their way to be incompatible with Proton...
Direct link to the talk, with timestamp: https://youtu.be/j5W5ErEMnvM?t=21385
Some other benefits of older GPUs:
Use as a dedicated GPU for encoding and decoding video. Post processing like frame interpolation or superresolution. Use for GPGPU workloads. Run additional monitors independently. Use for GPU passthrough to virtual machines. Use as a backup GPU for troubleshooting. Use for test code without breaking the main GPU
If only AMD did this.
AMD supports Mesa precisely so that enthusiasts can do this.
Nvidia doesn't, and their Vulkan stack underperforms on Linux quite significantly.
Except VFX and Hollywood have no qualms using proprietary drivers.
Even proprietary drivers can use Mesa. Most of Mesa is MIT-licensed.
It's just that Nvidia doesn't care much about Linux, and -as always- Nvidia ignores what everyone else is doing and does their own thing. Sometimes doing their own thing works out really well in the short run, but -long term- they always fall behind.
> It's just that Nvidia doesn't care much about Linux
Nvidia cares a lot about Linux. Just... on their terms.
Is there financial incentive for them to bother?
One thing they should've learned from Nvidia is that it's really worth it for them to invest making their devices function as broadly as possible. Crypto and AI waves both benefited Nvidia much more than AMD partly due to their devices being more universally usable.
That "partly" was worth hundreds of billions of dollars but AMD were cheap/shortsighted enough to hire a few dedicated engineers.
Trillions.
Perhaps not but one of the great things about Nvidia is their latest drivers often support very old cards.
They're in the process of ending support for everything below the RTX 2000 series -
https://windowsforum.com/news/nvidia-ends-feature-support-fo...
Perhaps René, the maintainer of T/2 Linux, will end up vibe coding an NVIDIA driver one of these days. He’s been reverse engineering and vibe coding a bunch of drivers for old graphics lately and live streaming everything.
Maybe there’s a prestige angle to somewhat shame them into keeping up with nvidia. Though my sense is that they’re barely peers anymore in this space. :(
Honestly, the folks writing inference drivers for Llama.cpp / GGML would sort of benefit from better compiler work like this.
In general, Valve's work has been exceptional and supplementing AMD's own ROCM/OpenCL and Vulkan team, they've gotten a lot of defaults right and people should work together with Valve to improve support for their chips.
I wonder if some of these could carry over for LLM inference. It will be nice to turn more ewaste GPUs into capable processing units for LLM.
Most of these older cards are lacking the physical hardware for fp8 or other lower precisions that most quantized models use. Or the memory to run models at higher precision.
I have a Radeon RX 6900 XT (16 GB VRAM, originally released in 2021), and it's possible to run some lightweight models to have okayish performance and quality of output, but nothing I've tried has come anywhere close to the quality even of the models I can use for free from OpenCode Zen or the free tier of Openrouter. If you want to keep everything local on the same card I have, it requires putting up with a model that's noticeably worse in virtually every metric than what you can get for free elsewhere, and the GPUs this article are talking about are three times as old as mine.
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
This kinda highlights the level of debt the AI companies are in, and will continue to be in, offering anything for free.
How long is this runway?
I have no clue, and I agree that it does not seem sustainable. Either someone needs to find a magic solution to making it a lot cheaper, or a lot of companies are going to need a lot of money from somewhere that isn't clear.
With an extra 8gb of vram you could run qwen 3.8 27b pretty comfortably, which isn't quite as good as frontier models but definitely on par with free models on openrouter and whatnot.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
Memory is the bottleneck along with the lack of FP4/FP8 capability at the hardware level.