• huppakee
    link
    fedilink
    arrow-up
    3
    ·
    16 hours ago

    I heard something like OpenAI (or Anthropic?) tweaked one of their AI’s to do well on a certain benchmark test and while they did have the highest score they couldn’t publish the model because it used like $20,000 in electricity alone to make sure it’s answer was correct.

        • davidgro
          link
          fedilink
          arrow-up
          5
          ·
          12 hours ago

          For context, this is a real question that stumped a lot of LLMs:

          “I want to wash my car. The car wash is 100 meters away. Should I walk or drive?”

          Of course it’s famous now, so they have added specific exceptions for it (or just because it’s famous they’ve scraped the solution)

          • [deleted]@piefed.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            11 hours ago

            It is also a great example of how low people’s expectations are of tech they have spent around a trillion dollars on to get something that is about as good as asking a random person.

            “People get the puzzle wrong!”

            We don’t spend a trillion dollars for someone to fail a simple logic puzzle.