I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.

These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):

  • Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
  • If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???


edit

Thanks to all who answered.

I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.

Asking to get arguments explained, I got more arguments instead.

  • NoLemurs@lemmy.world
    link
    fedilink
    arrow-up
    13
    arrow-down
    1
    ·
    23 hours ago

    It is absolutely possible to run 100% locally, but in practice, at the high end, only quantized models. The full size top end models require hundreds of GB of video ram, and while you can buy that, it’s stupidly expensive. Quantized models can often perform nearly as well with a small fraction of the ram, but they do sacrifice a little in precision.

    These models (well the good ones) ultimately all trace their origins to what you’d likely consider “stolen” data. Whether that’s ethical is debatable. If you’re in the “information should be free” camp, there may not be an issue here.

    As for the power/environmental impact, for what they do LLMs are actually very low impact per-request. If you’re concerned about your personal AI power use, then I hope you never fly in an airplane, and minimize your driving because those are much bigger issues.

    It’s the scale of use that makes AI an environmental problem, and that’s a question about corporate use of AI, not personal use of AI.

    • pemptago@lemmy.ml
      link
      fedilink
      English
      arrow-up
      1
      ·
      16 hours ago

      As for the power/environmental impact, for what they do LLMs are actually very low impact per-request.

      Worth noting that a request is often dozens of requests now that there’s “reasoning,” even for search. As I understand it, a model will take a question , figure out the context (one request), reform the question so it yields better results (another request), if it’s doing a web search there’s requests for each result, another to compare, another to check if it answers the original request, if not it loops and does it all over again. So one request is easily, and often, dozens of requests. This is one way Ai companies can say to investors, “see, look at how much usage has increased.”

      Also, we need to factor in the power to scrape and train each of those models, build the datacenters which is near impossible as these companies are not transparent about it and actively try to obstruct investigations into it. Then there’s the redundancy of all these different companies competing and doing roughly the same thing at the same time, as fast as they can, so it’s orders of magnitude inefficient energy consuming before it gets its first user prompt.

      Comparing it to other assaults on the environment is not only hard to do, but a case of “the worse negates the bad” fallacy.

      • NoLemurs@lemmy.world
        link
        fedilink
        arrow-up
        1
        ·
        15 hours ago

        I do see what you’re saying. You can account for all of these factors and it still turns out that, largely, individual LLM use just doesn’t use that much power compared to most things people do day to day. Inference is so cheap that even dozens of requests don’t amount to much. I could look up and give you a bunch of numbers, but I don’t think that’s likely to convince anyone who doesn’t do the research themselves. It’s so easy to come up with sources that say what you want. I’d encourage you to actually look into this yourself.

        Training costs are higher, but you train once and use repeatedly. Right now, total training costs are stupidly high, but that’s because we’ve got an arms race between the frontier labs to spend as much money and compute as they can for truly marginal gains in quality. The solution to that problem isn’t for individuals to stop using AI, it’s to stop those assholes from wasting so much power.

        Individual LLM use is so cheap, that it really isn’t worth wasting people’s energies thinking about limiting that. Instead of being distracted by attempts to make this an issue of personal responsibility, we should be focusing on what will actually make a difference. We should be focused on supporting policies that lead to systemic change. A carbon tax would change corporate behavior right quick, and not just for AI companies.

    • HubertManne@piefed.social
      link
      fedilink
      English
      arrow-up
      2
      ·
      20 hours ago

      The power impact was something that in the early days worried me. Looking into it I agree with you to some degree. Like using it instead of a search engine I think is by and large a wash. One prompt will likely take more energy but will give you information that likely would have required searching several times modifying the words and jumping between sites which are rendering all sorts of things. Heck If I booted into a command line and connected to an llm Im almost sure it would be significantly less energy. If you chat for entertainment instead of streaming vidoe also lower energy use. Now I think one thing is in making things. It lets people who otherwise couldn’t make pictures and videos and code. In the large majority of cases what is made is going to be disposed even for folks that eventually make something they care to keep around or use. While using software to do the same uses a lot of energy the only people doing it generally where making long lasting things for projects or such. So that is where I question it. Still I will have it make a picture to use in an rpg or such.

      • NoLemurs@lemmy.world
        link
        fedilink
        arrow-up
        1
        ·
        edit-2
        16 hours ago

        I’m going to simplify a little here, so don’t take this completely at face value.

        Models are, quite literally, long series of numbers (called weights). A model might store the weights in 16 bit numbers (that is, 16 binary digits). The size of the model (and how much memory it needs) is determined by how many weights there are, and how many bits each weight takes. You can take a 16 bit model and rework it to use 8 bit, or even 4 bit numbers. The result intuitively behaves a lot like the same model, but with less precision to the weights. That makes the model take way less space in ram, but also makes it more likely for concepts (encoded in the weights) to overlap, which impacts model quality. Often the effect is that fine distinctions get lost.

      • partial_accumen@lemmy.world
        link
        fedilink
        arrow-up
        1
        ·
        17 hours ago

        Kind of like “compressed”. It takes longer/more effort to run them, to produce the same result as a the same model that has not been quantized, where that non-quantized version would consume significantly more RAM but produce the result faster. You would typically only run the quantized model when you’re starved for RAM, which most of us are running LLMs locally.

        Think like zipping a file with file compression. It takes less space, but has to be unzipped for you to have usable files again.