This collaborative space allows users to contribute additional information, tips, and insights to enhance the original deal post. Feel free to share your knowledge and help fellow shoppers make informed decisions.
This is one 16GB, bandwidth around mid 400GB, next step up is 5070 TI at close to 800GB/s memory bandwidth which will make a huge difference in token generation speed also. It's also $829-$900 when a deal is on so depending on how much you wanna jump up to
$330 8GB -> $550 16GB 5060 TI. ->. $900 ~80% speed boost 5070 TI
based on memory bandwidth.
Our community has rated this post as helpful. If you agree, why not thank desynergy
Quote
from NeedBargain
:
Can i run a local LLM with this ? Only 8GB ?
We have a few AI development boxes at our office that came with a 8GB 5060 or a 5060 Ti with 16GB. The 16GB ASUS 5060 Ti's does circles around the 8GB version. We primarily deal with Qwen and Gemma (and a couple others you havent heard of). The 8GB version did Qwen 3.5 9B decently, but the 9B version is "dumb". It makes a lot more mistakes, and a lot of times errors out (lots of JSON errors). Qwen3.5 35B A3B on the 16GB ASUS card is a lot smarter, but takes longer to process. These 5060 Ti's are more for the 14B models, but again, 14B still made mistakes, just not as near as the 9B version did. It very much prefers the 32GB 5090 system ($4k video card), but I would say 16GB would be the bare minimum to run a LLM that isnt "dumb". We couldnt get the 35B to do anything on the 8GB systems, it would think for a couple minutes and say "f this". lol
They are getting them to work on weaker hardware, but the advancement is outpacing it by a wide margin, so while they're forking out better AI, getting them to work on lower end tech is a lot slower to be rolled out. If you do get a 8GB version, I would stick with the older models that have been tweaked to it's fullest, like Qwen 2.5. It doesnt have all the bells and whistles like the current versions, but they run better now on weaker tech.
Note: I'm just a tech that supports these systems. When it comes to development, I have no clue on any of that. I just know enough to troubleshoot and get the systems back operating. Also these lower end cards were purchased on purpose for the study/development they're working on. The actual servers have 4 - 6 headless nvidia cards in them that whoops all of these video cards butts, and run at least a 70B or higher LLMs. If you want to do higher end models on a budget, consider a Mac Studio. While nvidia is now putting Apple to shame, the Mac Studio still can do higher end models on a much slower basis.
Last edited by desynergy August 2, 2026 at 09:30 AM.
4
3
3
Like
Helpful
Funny
Not helpful
Sign up for a Slickdeals account to remove this ad.
We have a few AI development boxes at our office that came with a 8GB 5060 or a 5060 Ti with 16GB. The 16GB ASUS 5060 Ti's does circles around the 8GB version. We primarily deal with Qwen and Gemma (and a couple others you havent heard of). The 8GB version did Qwen 3.5 9B decently, but the 9B version is "dumb". It makes a lot more mistakes, and a lot of times errors out (lots of JSON errors). Qwen3.5 35B A3B on the 16GB ASUS card is a lot smarter, but takes longer to process. These 5060 Ti's are more for the 14B models, but again, 14B still made mistakes, just not as near as the 9B version did. It very much prefers the 32GB 5090 system ($4k video card), but I would say 16GB would be the bare minimum to run a LLM that isnt "dumb". We couldnt get the 35B to do anything on the 8GB systems, it would think for a couple minutes and say "f this". lol
They are getting them to work on weaker hardware, but the advancement is outpacing it by a wide margin, so while they're forking out better AI, getting them to work on lower end tech is a lot slower to be rolled out. If you do get a 8GB version, I would stick with the older models that have been tweaked to it's fullest, like Qwen 2.5. It doesnt have all the bells and whistles like the current versions, but they run better now on weaker tech.
Note: I'm just a tech that supports these systems. When it comes to development, I have no clue on any of that. I just know enough to troubleshoot and get the systems back operating. Also these lower end cards were purchased on purpose for the study/development they're working on. The actual servers have 4 - 6 headless nvidia cards in them that whoops all of these video cards butts, and run at least a 70B or higher LLMs. If you want to do higher end models on a budget, consider a Mac Studio. While nvidia is now putting Apple to shame, the Mac Studio still can do higher end models on a much slower basis.
I'd say it depends on what you're doing with it. I have qwen 3.5 9B running on a 2070 super 8gb and it's a little slow but it does what I want pretty well.
I have it doing some classification and data revision and it works well
IMHO avoid 8GB VRAM video cards! I even have a 10GB card from AMD (RX6700) and the models it can run are not very accurate at all. If you are going to do any meaning full work a better route will be go with an old 24GB VRAM models like RTX4090 or RX7900XTX; however, they are in a different class of GPUs. My experience with RX7900XTX is mostly positive.
Yes I don't like 8GB for AI. I tried a 12GB and was able to run LTX video generation (slow as hell) but it worked without messing around my setup that I run against the 16GB or 32GB VRAM Machines.
I can get LTX to run on 8GB if I use diff quantized files but then it means I have to run diff setup across diff VRAM machines and that's no bueno.
Look on the bright side, AMD is down hard today after ER. Maybe the AI bubble is finally gonna burst and make everyone poor and can't buy GPU even at lower prices.
8GB for some LLM fun or chain a couple together but then 2 8GB is like $400 so you are close to $530 for 16GB card that be more useful than 2 x 8GB. cards.
Quote
from nxh786
:
IMHO avoid 8GB VRAM video cards! I even have a 10GB card from AMD (RX6700) and the models it can run are not very accurate at all. If you are going to do any meaning full work a better route will be go with an old 24GB VRAM models like RTX4090 or RX7900XTX; however, they are in a different class of GPUs. My experience with RX7900XTX is mostly positive.
Join The Conversation
Share your experience with the Slickdeals community
Share information with the community. Please follow our Community Guidelines and be kind!
32 Comments
Sign up for a Slickdeals account to remove this ad.
Our community has rated this post as helpful. If you agree, why not thank xiaowang
Yes, but it won't do anything close to a large model. Something more like a 4 billion parameter model I think.
Our community has rated this post as helpful. If you agree, why not thank Elon69
Get at least 16GB so you have some flexibility.
https://slickdeals.net/f/19832085-microcenter-msi-nvidia-geforce-rtx-5060-ti-ventus-2x-black-plus-overclocked-dual-fan-16gb-gddr7-pcie-5-0-graphics-card-549-99?v=1&src=Site
This is one 16GB, bandwidth around mid 400GB, next step up is 5070 TI at close to 800GB/s memory bandwidth which will make a huge difference in token generation speed also. It's also $829-$900 when a deal is on so depending on how much you wanna jump up to
$330 8GB -> $550 16GB 5060 TI. ->. $900 ~80% speed boost 5070 TI
based on memory bandwidth.
Our community has rated this post as helpful. If you agree, why not thank desynergy
They are getting them to work on weaker hardware, but the advancement is outpacing it by a wide margin, so while they're forking out better AI, getting them to work on lower end tech is a lot slower to be rolled out. If you do get a 8GB version, I would stick with the older models that have been tweaked to it's fullest, like Qwen 2.5. It doesnt have all the bells and whistles like the current versions, but they run better now on weaker tech.
Note: I'm just a tech that supports these systems. When it comes to development, I have no clue on any of that. I just know enough to troubleshoot and get the systems back operating. Also these lower end cards were purchased on purpose for the study/development they're working on. The actual servers have 4 - 6 headless nvidia cards in them that whoops all of these video cards butts, and run at least a 70B or higher LLMs. If you want to do higher end models on a budget, consider a Mac Studio. While nvidia is now putting Apple to shame, the Mac Studio still can do higher end models on a much slower basis.
Sign up for a Slickdeals account to remove this ad.
https://www.newegg.com/gigabyte-g...6814932812
We have a few AI development boxes at our office that came with a 8GB 5060 or a 5060 Ti with 16GB. The 16GB ASUS 5060 Ti's does circles around the 8GB version. We primarily deal with Qwen and Gemma (and a couple others you havent heard of). The 8GB version did Qwen 3.5 9B decently, but the 9B version is "dumb". It makes a lot more mistakes, and a lot of times errors out (lots of JSON errors). Qwen3.5 35B A3B on the 16GB ASUS card is a lot smarter, but takes longer to process. These 5060 Ti's are more for the 14B models, but again, 14B still made mistakes, just not as near as the 9B version did. It very much prefers the 32GB 5090 system ($4k video card), but I would say 16GB would be the bare minimum to run a LLM that isnt "dumb". We couldnt get the 35B to do anything on the 8GB systems, it would think for a couple minutes and say "f this". lol
They are getting them to work on weaker hardware, but the advancement is outpacing it by a wide margin, so while they're forking out better AI, getting them to work on lower end tech is a lot slower to be rolled out. If you do get a 8GB version, I would stick with the older models that have been tweaked to it's fullest, like Qwen 2.5. It doesnt have all the bells and whistles like the current versions, but they run better now on weaker tech.
Note: I'm just a tech that supports these systems. When it comes to development, I have no clue on any of that. I just know enough to troubleshoot and get the systems back operating. Also these lower end cards were purchased on purpose for the study/development they're working on. The actual servers have 4 - 6 headless nvidia cards in them that whoops all of these video cards butts, and run at least a 70B or higher LLMs. If you want to do higher end models on a budget, consider a Mac Studio. While nvidia is now putting Apple to shame, the Mac Studio still can do higher end models on a much slower basis.
I have it doing some classification and data revision and it works well
I can get LTX to run on 8GB if I use diff quantized files but then it means I have to run diff setup across diff VRAM machines and that's no bueno.
Look on the bright side, AMD is down hard today after ER. Maybe the AI bubble is finally gonna burst and make everyone poor and can't buy GPU even at lower prices.
8GB for some LLM fun or chain a couple together but then 2 8GB is like $400 so you are close to $530 for 16GB card that be more useful than 2 x 8GB. cards.
Sign up for a Slickdeals account to remove this ad.
Join The Conversation
Share your experience with the Slickdeals community
Share information with the community. Please follow our Community Guidelines and be kind!