• DarkCloud@lemmy.world
    link
    fedilink
    English
    arrow-up
    11
    ·
    edit-2
    2 days ago

    Definitely true of data that needs to be tagged. Less true for text only data, which relies more on Markov chains.

    Interestingly the article doesn’t make this point, instead claiming workers in the global south are “feeding the data in” and doing maintenance work daily. So perhaps it’s on about moderation?

    • searabbit@piefed.social
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 days ago

      I only skimmed the article, but I’m familiar with this area of work. It’s the fine tuning phase of llms that require human “graders” in a sense to tell the model what is a helpful vs unhelpful response and elevate it from a literal autocomplete tool to a chatbot. The biggest platform for this kind of work is DataAnnotation if you want to look it up. It’s not just random people in the global south but also professionals training the model on their niche expertise (e.g., lawyers, accountants, engineers). The other type of work that has been going on even pre-llm chatbot era is, as you said, moderation. You moderate the chatbot for “bad” answers, correct them, and often that means impersonating the bot to the end user. I believe this is the main driver behind llm progress lately; literally humans teaching AI how to appear more human.

      • DarkCloud@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 days ago

        That’s funny because what’s human can change from culture to culture. Still globalization has been changing that… Making a more universal culture via the Internet.

  • pikl@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    2 days ago

    Same as every other technology. So long as you don’t look down you won’t notice the bony slave backs you’re standing on.