I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
edit
Thanks to all who answered.
I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.
Asking to get arguments explained, I got more arguments instead.


You need some decent hardware to do it, but yes one can absolutely run capable local models. The 7-8B models run fast and are probably decent for light coding, but they make piss poor chat bots. 27-36B models get much better outputs but take more hardware, and still not great in a chat interface (but far better than 8B). They run slow enough on my MacBook that it’s just a toy for fun, despite unified memory. I aspire to having a local AI server then can run models in the 100B-350B range. It’s a pricy bit of hardware, but I can leverage it to do all the things one might want to do but avoid for data privacy purposes. I try to avoid oversharing, but when I look at what ChatGPT knows about me, I’m giving way more away than I want to. I look forward to the day I don’t have to use it at all.
Ethically is a slightly different story; there are a number of claims people like to make:
Plagiarism: not in any significant sense. Like any of us might borrow a turn of phrase, so might AI, but any significant body of text — even a paragraph — is very unlikely to match exactly to an ingested work. That’s why they need to ingest so much different source material — so no one author or text dominates the output.
Stealing: Mostly yes. I think there is a significant difference between ingesting all of Reddit (where people including myself were sharing their words with the world) and feeding in copyrighted works of published authors without permission. I think when you turn around and sell access to a model built on material you didn’t pay for, that’s wrong. I do think open models available to everyone are far less evil because they take in information they don’t own and give back a tool they don’t own, having contributed significant money in the form of training.
Infrastructure: Yes but… that’s how everyone does it from Walmart to sports stadiums. Don’t single out data centers when the entire system is corrupt. Companies should pay for transmission lines and capacity they expect to draw. All companies, not just AI.
Environment: Yes but… the amount of water used sounds high on a global basis, but it’s utterly dwarfed by agriculture and some other industries. Now some say, yeah but we have to eat, but do we have to eat almonds specifically? Beef? Do we have to grow food in deserts? Data center water usage looks big in raw numbers, but it’s not in percentage terms. That being said, it can be a huge problem locally where a community might have a sustainable water supply but a big data center tips the scales.
The latter two are IT industry issues that are driven by AI but aren’t going away when and if the bubble pops. Imagine the privacy nightmare that is coming. In fact, I’m not convinced that AI isn’t just a cover for unprecedented data collection.