• MalReynolds@slrpnk.net
    link
    fedilink
    English
    arrow-up
    330
    ·
    14 days ago

    Trying to find it funny, but in 2026 it’s way too close to the bone and just makes me sad.

    • jafra@slrpnk.net
      link
      fedilink
      arrow-up
      20
      ·
      edit-2
      14 days ago

      Yeah. I first thought it’s about jpeg, then i read ai and i wasnt sure if its worrying stupidity reporting or bad satire. Edit: i thought '92 i mean

    • AmyAye@nord.pub
      link
      fedilink
      English
      arrow-up
      8
      ·
      14 days ago

      Yeah, like how they are destroying rare books.

      I mean, you don’t get that back just by asking the content theft machine.

    • Maeve @lemmygrad.ml
      link
      fedilink
      arrow-up
      6
      ·
      edit-2
      14 days ago

      It’s what Apple and MS basically do by uploading all your images to iCloud and showing the name in your images folder.

      Edit to ask: is this what the human brain does when retrieving a memory?

  • DaddleDew@lemmy.world
    link
    fedilink
    arrow-up
    190
    ·
    edit-2
    14 days ago

    Amateur. I can compress entire seasons of a TV series to a few bytes. All I have to do is type its title in Netflix and then BAM, gigabytes of video come out.

  • OldGrayDog@fedinsfw.app
    link
    fedilink
    English
    arrow-up
    58
    ·
    14 days ago

    I’ve read that the Trump administration is hiring him to archive all of the Epstein files using that format.

  • ShellMonkey@piefed.socdojo.com
    link
    fedilink
    English
    arrow-up
    42
    ·
    14 days ago

    On a similar note, I saw a story a bit back of someone saving input tokens by feeding the bot an image of a wall of text rather than the text itself and having it read the image via OCR.

    Satire and reality are too hard too distinguish these days.

    • lad@programming.dev
      link
      fedilink
      English
      arrow-up
      19
      ·
      edit-2
      14 days ago

      That’s nice, albeit I want to point out for anyone wondering that this is only conjectured and not guaranteed:

      One of the properties that π is conjectured to have is that it is normal, which is to say that its digits are all distributed evenly, with the implication that it is a disjunctive sequence, meaning that all possible finite sequences of digits will be present somewhere in it.

      There is no guarantee for any specific sequence to appear in π, but for short chunks chances are better (it’s not really a probability, but it’s simpler to say and I can’t explain in details anyway). That’s because (from wiki):

      It is widely believed that the (computable) numbers √2, π, and e are normal, but a proof remains elusive.

  • whereitsat@lemmy.zip
    link
    fedilink
    arrow-up
    25
    ·
    13 days ago

    brilliant satire that critiques all of magazine journalism.

    i love going to [insert publication here] and reading another article about ‘so and so is ready for their next chapter in life.’ the so and so always an uninteresting, overly wealthy fuckwit that hasn’t accomplished anything other than going to college and having a wealthy parent.

  • TrickDacy@lemmy.world
    link
    fedilink
    arrow-up
    25
    ·
    14 days ago

    a typical jpeg of 20 Mb

    Uh what. Literally should be the highest fucking possible quality from a $10K camera if it’s that big. My raw images aren’t even that big usually.

    • PieMePlenty@lemmy.world
      link
      fedilink
      arrow-up
      18
      ·
      edit-2
      14 days ago

      Its not typical, but you can get 20 Mb+ jpegs out of an entry level 18MP dlsr. Especially if there’s lots of color and at like 5500x3300 resolutions and created with 100% quality preset.
      I checked my immich and I have some (and larger), but yeah, not exactly typical.

      • TrickDacy@lemmy.world
        link
        fedilink
        arrow-up
        2
        ·
        14 days ago

        I am not sure I’ve ever used the 100 quality setting on jpegs. Many years ago I experimented with that a lot and decided that anything over 90 was not different to my eye but the file size was much bigger, relatively speaking. So yeah I suppose if you did use 100, a 20 MB jpeg image is not hard to reach.

    • bstix@feddit.dk
      link
      fedilink
      arrow-up
      15
      ·
      14 days ago

      Raw images is where you go wrong. You need at least 10 mb of metadata tags to achieve professional levels of file sizes. How can you even look at a picture without having a full description of all your childhood memories that lead you to take this beautiful picture of yesterdays mac’'n’cheese dinner. This is why we need more data centers.

  • M1k3y@discuss.tchncs.de
    link
    fedilink
    arrow-up
    20
    ·
    14 days ago

    The sad thing is that this has been possible for decades using convolutional autoencoders, but with LLMs we forgot that AI architectures other than transformers still exist.

    • CanadaPlus@lemmy.sdf.org
      link
      fedilink
      arrow-up
      6
      ·
      edit-2
      13 days ago

      Yeah, it’s really awful. With any luck, AI winter will follow AI summer, like usual, and the serious people can come out again. Although, aren’t CNNs more of a this century thing? I guess two decades is still decades…

      IIRC autoencoders actually produce the same image to within our ability to notice, as well.

    • Axolotl@feddit.it
      link
      fedilink
      arrow-up
      7
      ·
      14 days ago

      Before you consider using the project, I highly recommend you read and understand the last paragraph of the LICENSE file. If you are seriously considering using it, it will become important.

      This is so funny

      For whoever is wondering, the license is MIT

      • boonhet@sopuli.xyz
        link
        fedilink
        arrow-up
        5
        ·
        14 days ago

        And also

        Ironically, aside from the obvious use of LLMs for the actuall project, everything was written by hand. I wanted to use this as a learning opportunity about git filters, so everything you see is my fault.

  • AItoothbrush@lemmy.zip
    link
    fedilink
    English
    arrow-up
    16
    ·
    14 days ago

    The fuuuucking annoying part is that these weight trained models would be perfect for translation models, compression, etc. An llm is already kind of a really efficient lossy compressor but you could actually make it lossless and an actual compressor if used properly. But instead people are literally telling llms to translate instead of training models that are for translating. The technology isnt the problem itself, its the industry and capitalism.

      • lad@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        ·
        14 days ago

        I think, they don’t mean lossless compression with LLM, but with a neural network in general. I think it might be possible, but I’m not sure

      • AItoothbrush@lemmy.zip
        link
        fedilink
        English
        arrow-up
        1
        ·
        14 days ago

        Yeah you can? A neural network that only relies on wheights and doesnt use random numbers will always give the same output for the same input.

    • ChaoticNeutralCzech@feddit.org
      link
      fedilink
      English
      arrow-up
      5
      ·
      edit-2
      14 days ago

      There are not many places where 95% compression of UTF-8 plain text could outweigh needing a 30GB model in memory and a significant fraction of current LLM inference cost to decompress it − and good luck convincing librarians to adopt it.

      • Natanael@infosec.pub
        link
        fedilink
        arrow-up
        2
        ·
        edit-2
        14 days ago

        There’s a 300 MB library for audio compression using it. If you have large audio libraries it could eventually become worth the tradeoff.

        https://huggingface.co/facebook/encodec_32khz

        The image versions are probably more useful though.

        The important part for these codes is that it has pre-LLM functions for quality metrics to judge if the output is close enough to indistinguishable (preserves detail, doesn’t add any).

        Although there is also variants deriving a neural net from the media to recreate it from the smaller model.