One-slot, low-profile Nvidia RTX 3060 12 GB with two monitor outputs breaks cover at Newegg for $496 — bus-powered model looking for a use case in local LLM wor

One-slot, low-profile Nvidia RTX 3060 12 GB with two monitor outputs breaks cover at Newegg for $496 — bus-powered model looking for a use case in local LLM wor

Bruno Ferreira Social Links Navigation Contributor Bruno Ferreira is a contributing writer for Tom's Hardware. He has decades of experience with PC hardware and assorted sundries, alongside a career as a developer. He's obsessed with detail and has a tendency to ramble on the topics he loves. When not doing that, he's usually playing games, or at live music shows and festivals.

me-three Admin said: Attentive readers might surmise that one (or more) of these would be good candidates for an entry-level local LLM rig. That's precisely how SRhonyra is pitching the card Yep, multi-GPU setup is one way to break the 24GB VRAM barrier of individual GPUs. Below, the YT'er DigitalSpaceport is using a single EPYC server board, modded with PCIe 4 risers, to accommodate quad-3090's, mounted in an open-frame setup. LLMs would have layers split over across all cards via CUDA device offloading via llama.cpp. Or using vLLM to serve with tensor parallelism across the GPUs. Admin said: Even then, it's likely that people interested in these cards are looking to use them either as secondary GPUs, or use more than one in the same box to be able to virtually pool their VRAM and use larger models than you'd otherwise be able to. Yep, the power draw cap is a deal breaker, along with the cooling issues. You're better off using standard (used) 3060s for a poor-man's version of what DigitalSpaceport built. 4x12GB=48GB is sufficient for mid-sized models, but even full-power 3060 has too weak compute to run such sized models. AI power users tend to opt for 3090s in either 2/3/4 setups. This 3060's high asking price does highlight a real demand imbalance for local AI use. It's the likely reason why 16GB GPUs are priced disproportionately high–especially 5060 Ti 16GB–because AI users are scarfing them up to use in multi-GPU rigs. This demand won't abate, but will only rise with time, as local AI use is trending upward, away from API use. Before the latest price craze, a used 3090 commands $1K, and a 5060 Ti 16GB goes for $500. On a bang/GB basis a quad-5060Ti setup is more cost effective, albeit with a lower VRAM ceiling. Some people think that if the AI bubble will only burst, then GPU pricing will return to normal, but the opposite will be true, as AI use will shift onto local models, and GPU pricing will jump even higher. The likely outlook is that Nvidia/AMD/Intel will expand AI-specific accelerators to encompass the prosumer and eventually consumer market, to meet the burgeoning demand. Intel at least will likely forgo the (gaming) GPU market altogether, and the same economic rationale exists for AMD or Nvidia to abandon or "freeze" gaming GPUs to serve AI-GPUs. We'll see what companies have in store for 2027. The local AI demand is definitely there. PC gamers may well be left out in the cold. So7tqRSZ0s8 View: https://www.youtube.com/watch?v=So7tqRSZ0s8 Reply

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment