• 3 Posts
  • 415 Comments
Joined 2 years ago
cake
Cake day: March 22nd, 2024

help-circle


  • Duct it!

    I have a 400W 3090 with zero case fans.

    Its always cool, because I bought like $10 of weather stripping to “seal” its intakes against the edge of an SFF case. So it’s always sucking in ambient air:

    You can also severely undervolt a 5090 to like 250-300W, with very little performance impact. Honestly its stock speeds are kind of crazy.

    …But again, you aren’t gaining much over a 4090. A 4090 + DDR4 threadripper would be way faster than a 5090 + occulink, for many reasons. And TR CPUs are pretty reasonably priced compared to GPUs these days.


    But even if you do go 5090, I’d highly recommend finding some way to shove it in the case and duct some air into its intakes. Its going to be way faster on a PCIe slot.

    You could even get a riser and mount it somewhere else in the case, theoretically.


  • First of all, I mean zero offense with any purchase decision. A 5090 is very good.

    …But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

    Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

    Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090, where a 3090 or 4090 will get the job done.

    Even if pure speed is your primary concern, you can tune a sparse 120B model (like Laguna) to run almost as fast as Qwen 27B on a 5090, and get at-least-good results.


    It’s more finicky and involved, though. For sure.

    Running an LLM on a 5090 is a task. Hybrid CPU + GPU inference is a hobby.



    • It’s just under 300B, trained at FP4; absolutely the perfect size for servers with 128GB-192GB CPU RAM to spare.

    • Its fast. I’m getting 17 tokens/sec on a single RTX 3090 GPU, all experts offloaded to RAM; for a 300B model this smart, that’s crazy fast.

    • Its attention mechanism is cutting edge, good for long context without too much processing time.

    • The benchmarks for coding/agenic usage are absolutely bonkers, within margin of error of frontier models or Deepseek Pro in some cases: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

    • …Though I don’t put much stock in benchmarks. And I haven’t tested it enough to tell you if it lives up to that hype for specific use cases.

    • Deepseek also publishes its base model. That means I can “unfry” the model with a merge if I have to.


    I am afraid of the the model being “overfit” to coding and agenic stuff.

    For reference, my previous favorite model was Xiaomi MiMo V2.5 310B. It benched well, but it also feels “smart” outside of benchmarks, like in knowledge of trivia without tool usage/internet access or comprehension of weird questions.


  • [AIT] I know this isn’t everyone’s cup of tea, but I’m excited about Deepseek V4 flash. It’s (for me) the perfect size and architecture to self-host an LLM.

    My box (and brain) are chugging through a queue:

    • Figure out why my swap is going crazy, and how to ban processes from it [Done].

    • Figure out why Code OSS is unhappy [Partially Done].

    • Make an ik_llama.cpp iMatrix for Deepseek V4 [Done].

    • Figure out why quantization isn’t working [Done].

    • Make a test IQ2_KL/MXFP4_R8 quant to see how it does squeezed onto my box [in progress].

    • Test. Tune. Inevitably troubleshoot the dozen other things that go wrong. Figure out how much spare RAM that leaves me.

    • Make a higher quality IQ3_KT quantization. This will take all night on my CPU.

    • KLD test both of them vs the full precision, to quantify quantization loss. Likely an overnight task, too.

    • Try merging the new model release with the base model, 50/50, for a less “deep fried” model. imatrix, quant, test.

    The goal is to host it on a single RTX 3090, Ryzen 7000 with 128GB RAM, for anyone curious. Though I may try smaller models too, like Laguna S1.







  • One more thing.

    If you look at the artifacts on the poster, they look replicated. For instance, look at the little mark to the left of his face:

    That, and other hints, are sus.


    …Does that mean it’s AI? Maybe. Maybe they printed that way to give it a “rough” look, or maybe all the posters were in a stack when they got damaged. Or maybe someone asked ChatGPT or Gemini to tessellate this poster across the wall, which has the exact same artifacts:

    img

    Or maybe you can find the original poster JPEG if you keep digging.


    Also, the original image? 6000x4000 pixels, precisely.

    That the exact resolution of 24MP camera sensors (like the Canon R8 or R50), but it’s also something AI could output, too.


    I just don’t want you to misinterpret me: this could absolutely be AI. 100%.

    Could be real. I think its real.

    But you’ll never get an answer, and don’t believe anyone who tells you otherwise.

    I’m arguing, with all due respect, that your original question of “Is this AI” does not matter, because you can’t trust random posts from strangers anyway. You never could. I’m not trying to be impolite, but that’s intended to be a cold bucket of water in the face.


  • Eyeballing the photo is hopeless.

    If you want to check credibility, do it the old fashion way, before AI existed: Check the source.

    I’ll take a stab.


    OP is: https://old.reddit.com/user/RustySchmeckleford/

    Post history is hidden. And in their entire history, they only have two comments. Weird, but inconclusive.

    Dead end.


    No results on Tineye. The exact image is one day old on Google Images, but it is showing a few visually similar posts. For example:

    https://www.theguardian.com/us-news/2026/jan/31/a-week-on-the-block-where-alex-pretti-was-killed

    similar wall

    An original photo from The Guardian? Reasonable to assume that’s real. But that’s from Minneapolis.


    Alright… let’s look for clues. What’s NRG Arena?

    An arena in Houston, Texas, as it turns out. Found this on Facebook searching for that text:

    That’s definitely the poster. This is probably some real street corner in Houston. That concert was in 2025, so It’s reasonable to assume the poster survived into 2026.

    And that lets us narrow down the search. For instance, local Houston news posted this:

    https://cw39.com/news/local/houston-ice-agent-shooting/

    https://cw39.com/wp-content/uploads/sites/10/2026/07/salgado_araujo_memorial3.jpg

    https://cw39.com/wp-content/uploads/sites/10/2026/07/salgado_araujo_memorial3.jpg


    CNN posted the original photo the poster is based on. Again, quite credible:

    https://www.cnn.com/2026/07/07/us/houston-ice-shooting-death

    CNN Photo



    So… is the photo fake?

    Shrug.

    RustySchmeckleford could’ve taken a random 2025 photo, a picture of the poster, and AI’d it onto the wall. It’s impossible to tell.

    But all the components are real.

    I’d argue you’re asking the wrong question. If some rando in the internet posts a picture, AI or no AI you have no idea if it’s real or not. You have to ingest it with that in mind.

    To be blunt: you will never know if it’s AI or not.

    You could dig into the Reddit user or analyze it all you want, you could trace the source of the poster template or pixel peep for artifacts, but fact is this post has zero credibility. That’s true of anything from internet strangers.

    If you want credible, get info from credible source. It’s that simple. /u/RustySchmeckleford/ is not The Guardian.


  • Encoder-decoder language models and all sorts of stuff were used for translation and spellcheck, long before “LLM” was in anyone’s vocabulary. Embeddings models were used in documentation searches, in IDEs, and other places. Whenever you used any search engine, pre Sam Altman, you were likely hitting text models too.

    They worked alright.

    It was not an issue. No one hated them; they are simple tools with a specific function.

    I think people need to be careful of spilling (quite reasonable) hate of Tech Bro AI into the wider, older field of machine learning. In spite of the effort to conflate them, they aren’t the same thing.





  • JXL is working alright for me.

    As an example with a lot of dynamic range, here’s a JXL:

    JXL image should be here

    AVIF:

    AVIF image should be here

    Both render in my desktop and iPhone browsers, just fine. I bet at least one renders for you. And I made them from RAWs from a really old camera!

    The problem is, as you say… arbitrary lack of support. As an example, I can’t upload either file to Lemmy. Brand new social media software, and it doesnt’ recognize JXL or AVIF as valid image types, even though they should render just fine? Most image hosts wont take JXL either, hence I had to upload them to litterbox since catbox is down!

    An HEIF, on the other hand, has basically 0 support outside of Apple:

    HIF Image should be here

    All three of these render correctly on my phone, but only the top two do on other devices.