image description

An infographic titled “How To Write Alt Text” featuring a photo of a capybara. Parts of alt text are divided by color, including “identify who”, “expression”, “description”, “colour”, and “interesting features”. The finished description reads “A capybara looking relaxed in a hot spa. Yellow yuzu fruits are floating in the water, and one is balanced on the top of the capybara’s head.”

via https://www.perkins.org/resource/how-write-alt-text-and-image-descriptions-visually-impaired/

  • li10@feddit.uk
    link
    fedilink
    English
    arrow-up
    28
    arrow-down
    1
    ·
    2 years ago

    Bro I fucking love capybaras so much

    10/10 animal, fucking brilliant.

    • Viking_Hippie
      link
      fedilink
      arrow-up
      4
      ·
      2 years ago

      See also: self referential

      goes to dictionary entry for “recursion”

  • airbussy@lemmy.one
    link
    fedilink
    English
    arrow-up
    16
    arrow-down
    3
    ·
    2 years ago

    Potentially also useful for creating good prompts for AI image generators?

    • Daxtron2@startrek.website
      link
      fedilink
      arrow-up
      7
      ·
      2 years ago

      It’s essentially by-hand CLIP, that’s how the training data for CLIP came into being, it was descriptive text for images.

    • pennomideleted by creator
      link
      fedilink
      English
      arrow-up
      6
      ·
      edit-2
      2 months ago

      deleted by creator

    • 9488fcea02a9@sh.itjust.works
      link
      fedilink
      arrow-up
      6
      ·
      2 years ago

      Prompts are just the reverse of image recognition AI tagging stuff.

      Alt text is exactly the kind of tedious work that AI would be good at doing, but everyone in the fediverse seems to have a huge hate boner for ANYTHING AI…

      Fediverse: write a fucking essay every time you post an image… But make sure you waste time doing it manually, instead of using AI tools!!!

    • Blaster M
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 years ago

      If you have really detailed image tags, a model trained on them can make great outputs.

  • biptoot@lemmy.today
    link
    fedilink
    arrow-up
    14
    arrow-down
    1
    ·
    2 years ago

    This is excellent, very useful for continuing to make images accessible on the fediverse

  • bjornsno@lemm.ee
    link
    fedilink
    arrow-up
    1
    ·
    2 years ago

    Ignorant question: isn’t alt text primarily for visually impaired people? If so, what is the point of including info about color?

    • Arthur Besse@lemmy.mlOPM
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 years ago

      Color can provide useful context. For example, in the case of this image, imagine if in a thread about it there was some discussion of the ripeness of the yuzu fruit.

  • drislands
    link
    fedilink
    arrow-up
    1
    ·
    edit-2
    2 years ago

    EDIT: Turns out I totally misunderstood what this graphic was communicating. Thank you Starkstruck for patiently exposing it to me!


    ORIGINAL COMMENT:

    I don’t like this. It feels like a lot of extra work for little benefit.

    That being said I’d love to hear from someone for whom this is helpful. Happy to be wrong.

      • drislands
        link
        fedilink
        arrow-up
        1
        ·
        2 years ago

        How does color coding alt text help blind people?

        • Starkstruck
          link
          fedilink
          arrow-up
          2
          ·
          2 years ago

          Oh geeze you seem to have completely misunderstood the point of this graphic. You’re not supposed to color code your alt text, the color coding is a guide to correspond to the labels at the top. It’s teaching you how to write good alt text, what to include and such.

          • drislands
            link
            fedilink
            arrow-up
            2
            ·
            2 years ago

            OH. Yes, if that’s the case I absolutely misunderstood. Wow, the way you describe makes a LOT more sense. Doh 😖

            Thanks for helping me understand!

  • flashgnash@lemm.ee
    link
    fedilink
    arrow-up
    2
    arrow-down
    1
    ·
    2 years ago

    Is this not the kind of thing machine vision/language models would be really good at?