Replaced $40/month in AI API subscriptions with self-hosted Ollama + n8n

quickbitesdev@discuss.tchncs.de · 2 months ago

Replaced $40/month in AI API subscriptions with self-hosted Ollama + n8n

doodledup@lemmy.world · 2 months ago

Running a thousand watts and not running a thousand watts can be quiet a difference depending on where you live. And then consider buying all of the hardware. In many cases it’s probably cheaper to just pay $40 al month.

StripedMonkey@lemmy.zip · 2 months ago

That would be true worst case, but you’re never running inference 24/7. It’s no crazier than gaming in that regard.

fuckwit_mcbumcrumble@lemmy.dbzer0.com · 2 months ago

What whack ass setup so you think OP has? Dual 5090s? They’re running it on an i7.

T156@lemmy.world · edit-2 2 months ago

It’s also an 8 gigaparameter model. That’s pretty tiny, even if they use it heaps.

Semperverus@lemmy.world · edit-2 2 months ago

Do you think it runs at 1000w continuously? On any decent GPU, the responses are nearly instantaneous to maybe a few seconds of runtime at maybe max GPU consumption.

Compare that to playing a few hours of cyberpunk 2077 with raytracing and maxed out settings at 4k.

Don’t get me wrong, there’s a lot to hate about AI/LLMs, but running one locally without data harvesting engines is pretty minimal. The creation of the larger models is where the consumption primarily comes in, and then the data centers that run them are servicing millions of inquiries a minute making the concentration of consumption at a single point significantly higher (plus they retrain the model there on current and user-fed data, including prompts, whereas your computer hosting ollama would not.)