• taiyang
    link
    fedilink
    arrow-up
    61
    arrow-down
    1
    ·
    2 days ago

    Can confirm; I teach statistics and allow take home exams. I don’t even really need guardrails on the math-- if you cheat via LLM, your answer is almost always hilariously wrong.

    • Mika@piefed.ca
      link
      fedilink
      English
      arrow-up
      8
      ·
      2 days ago

      I’ve not used gpt for quite a while, but can’t you ask to python script the evaluation?

      • EnsignWashout@startrek.website
        link
        fedilink
        arrow-up
        26
        arrow-down
        1
        ·
        edit-2
        6 hours ago

        For simple math that could work, and as long as the question is close enough to an exact match with plenty of published examples to copy.

        A good rule of thumb is that the script it will come up with is about as likely to be correct as blindly taking the highest voted answer to the most similar question on Stack Overflow.

        If the question is simple and common, the odds are quite good. If the question is nuanced or rare, the odds of a correct result drop off aggressively.

        Edit: Your mileage may vary - providing an API to these LLMs that can do math correctly for them is pretty easy. Getting the LLM to consistently detect when to use that API is more challenging. But work-arounds exist.

        • gandalf_der_13te@feddit.org
          link
          fedilink
          arrow-up
          2
          ·
          23 hours ago

          cannot confirm. have used chatgpt many times to check my physics homework (after doing it myself first) and it was basically correct about 90% of the time.

          these were not standard exercises (i think)

          • EnsignWashout@startrek.website
            link
            fedilink
            arrow-up
            1
            ·
            21 hours ago

            Thanks. I think every data point helps people.

            I can’t say that particularly shocks me. I imagine that Chat GPT has probably swallowed dozens of Physics textbooks.

            The bigger part of the problem is knowing where we are on the human knowledge novelty / training data available curve, or rather, when we need to get off of the ride.

            Edit: Any chance your chosen AI has access to Wolfram Alpha as an MCP Tool? Because that would be very different, too.

            If it’s successful at math, there’s likely much more than a learning model in play.

            Which brings us back to the opacity problem. :(

            If the answer machine has been cleverly rigged (and weighted) to pass off math questions to more capable software, things can be great.

            But a person who just hears that AI can do math now can have a very bad time if they go blindly use a brute force pleasant answer machine on a math problem. :(