
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
Oscar Molnar explains that a cheap Tesla V100 SXM2 with 16GB HBM2 was sourced, as was an SXM2-to-PCIe adapter, and a PWM mod for the loud-as-a-lawnmower cooler, to complete this VRAM expansion for the hefty local LLMs project. Indeed, these GPUs do look cheap right now, as I can see them listed on eBay US for under $140 each, if you don’t mind buying from China.
As mentioned above, you can’t just get one of these Tesla V100 SXM2 cards with abundant VRAM and plug it into your PC. Molnar says they spent about $66 on an SXM2-to-PCIe adapter , also on eBay.
You might think that was enough. However, the PC and local LLMs enthusiast baulked at the noise of “the fan from hell,” which came as standard with the Tesla V100 SXM2. That shrieking cooler was measured outputting 82dB of noise. Molnar described it as “somewhere between a garbage disposal and a lawnmower.” This may be the most complicated tweak yet, but basically the existing fan wires just needed rerouting and plugging into the motherboard PWM fan header. You could also simply purchase a “2.54mm male to PH2.0 female jumper cable” for the task. Apparently, the fan only needs to run at 10% to keep the Tesla V100 under 50C at full load.
With the hardware all now fitted and finessed, Molnar had a 32GB VRAM system at their disposal – that’s a PC with RTX 4080 : 16GB VRAM, Ada architecture and Tesla V100: 16GB VRAM, Volta architecture . They note you can get Tesla V100s with 32GB of VRAM, but they are double the price.
$200 Nvidia AI GPU for servers hacked into a PCIe card with custom PCB and 3D-printed cooling
768GB of cheap Intel Optane DIMM memory sticks used to run 1-trillion-parameter LLM on a system with a single GPU
Fully-functional RTX 3070 16GB gets frankensteined into existence by harvesting dead PCBs and RX 6800 XT's VRAM chips
Getting the system to make use of this 32GB of total VRAM for LLMs wasn’t tricky, says the DIYer. They used NixOS with a legacy Nvidia driver that overlapped support for both Volta and Ada architectures. Testing a local LLM , they got a 27 billion parameter model running at 32 tokens per second, which they say is “fast enough for interactive use” and faster than most cloud API alternatives.
Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/pc-components/gpus/SPONSORED_LINK_URL
- https://www.tomshardware.com/pc-components/gpus/ai-enthusiast-adds-nvidia-tesla-v100-as-loud-as-a-lawnmower-to-gaming-pc-for-usd266-32gb-of-vram-rig-can-run-27-billion-parameter-model-at-32-tokens-per-second#main
- https://www.tomshardware.com/subscription
- Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
- Physicists turn particles in chaotic orbits into liquid computers — but this fluid hardware still trails memristor rivals
- China begins mass production of homegrown immersion chipmaking machines in major breakthrough, report claims — first DUV lithography units will be delivered thi
- Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
- Save $1,000 on Lenovo's over-the-top RTX 5090 gaming laptop — this 18-inch monster packs 64GB of RAM and up to a 440Hz refresh rate
Informational only. No financial advice. Do your own research.