• grue@lemmy.world
      link
      fedilink
      English
      arrow-up
      72
      ·
      9 days ago

      Because Swartz was merely a natural person, while Meta is an almighty corporation. Everybody knows only corporations deserve rights, duh!

    • Duamerthrax@lemmy.world
      link
      fedilink
      English
      arrow-up
      34
      ·
      9 days ago

      Because Swartz wanted a free and open web and the powers that be wanted control over every major social media website. This was around the same that moot showed up in Epstein’s emails and would later sell his site. All in the run up to the 2016 election. Now think about how /r/thedonald stayed on the site for so long.

      • Daftydux@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        4
        ·
        8 days ago

        r/donald was a py-op posing as social media. I doubt it did much other than get liberals mad on reddit but it was clearly a sign of things to come.

        Fox is the true evil.

    • John Richard@lemmy.world
      link
      fedilink
      English
      arrow-up
      11
      ·
      9 days ago

      Lol, doesn’t matter if you have expensive lawyers. What matters is that you have counsel that goes to church with the judge or plays golf at the country club with them, or knows people that do. Or you’re a Zionists genocide supporter and pedophile like Alan Dershowitz.

    • John Richard@lemmy.world
      link
      fedilink
      English
      arrow-up
      27
      ·
      9 days ago

      Fuck Congress and the DoJ as well. How the fuck can Reddit still have Section 230 protections after what spez did?

      • Unsealed9041@lemmy.ca
        link
        fedilink
        English
        arrow-up
        9
        ·
        9 days ago

        Democracy was always a mirage, it was only for rich white guys in the beginning, not the enslaved Africans and genocided natives. We’ve found ways around giving liberty to at least some portion of people continuously since.

      • JoeBigelow@lemmy.ca
        link
        fedilink
        English
        arrow-up
        8
        ·
        9 days ago

        More like we’ve known, but it’s been more tolerable until recently. We got frog boiled, and are just realizing we’re cooked.

  • Zephyr@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    43
    ·
    8 days ago

    Moral of the story never do anything as an individual by your name. Do it as a multi-billion dollar company with a battalion of lawyers and have fall guys. At minimum have an LLC controlled by a trust in someone else’s name doing anything actionable in court.

    Remember companies are people even though they can’t get arrested and giving large sums of money to politicians is free speech, not bribery or corruption.

    • UsoSaito@feddit.uk
      link
      fedilink
      English
      arrow-up
      5
      ·
      8 days ago

      And do it in a large enough quantity that it makes financially insensitive to do it at a large scale vs small

      • Zephyr@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        11
        ·
        8 days ago

        In short just add lots of people and many steps in the way. Killing someone is bad, releasing a product / service you know will kill lots of people is just a rounding error in business.

        • busted_Anoose@aussie.zone
          link
          fedilink
          English
          arrow-up
          2
          ·
          7 days ago

          one only has to look at the health insurance industry to see how deadly that is both for claimants and ceo’s alike.

  • helvetpuli@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    39
    ·
    9 days ago

    So scraping means any harvesting of data now?

    We need a new name for the fairly painful process of trying to tease meaning out of people’s unstructured HTML.

    In any case downloading a bunch of stuff from JSTOR was not scraping. And he absolutely had authorised access to that data. They took exception to the quantity, mainly.

    • tiramichu@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      10
      ·
      edit-2
      8 days ago

      Scraping has multiple meanings.

      Web scraping is a specific type of scraping, but data via APIs or even torrents could be considered a scrape, even if that data is nicely structured.

      The commonality between them is they all have the implication that:

      • the data harvesting is automated
      • the data you harvest is not owned by you, and you don’t have explicit permission to use it
      • the scope of what you harvest is broad and not targeted at retrieving specific limited pieces of data

      Any access patterns that broadly correspond to this could be considered scraping.

      • helvetpuli@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        2
        ·
        7 days ago

        Sure. Fine. It means all of that now.

        But what are we going to call the difficult thing that we have to do to coax, say, unstructured event listings into reasonably structured data?

        • tiramichu@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          1
          ·
          7 days ago

          If you want to refer specifically to web scraping then call it “web scraping” or “site scraping” or “HTML scraping”

          Not difficult to get your meaning across.

    • Nugscree@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      8 days ago

      Also that he was planning, or already was, sharing the data for free. They made an example out of him because he believed such data should be for everybody.

  • HubertManne@piefed.social
    link
    fedilink
    English
    arrow-up
    38
    ·
    9 days ago

    Aaron is definately on my wall of heroes. I think much of our leadership does not understand how despite there being billions of people on the planet they are not just fungible commodities. That we lose every time and stunt our growth. I by no means mean we should have more populationa as what we have is to much for the biosphere. Its about smartly using what we have than thinking you can just replace a person with another.

  • Uriel238 [all pronouns]@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    24
    ·
    9 days ago

    The anti-piracy efforts of the RIAA and MPAA (and the publishing houses and…) are still pretty robust, but that didn’t stop any of the big AI companies from using gigatons of copyrighted material as datasets to train their LLMs. It’s why when you ask them to generate an image featuring Winnie The Pooh, they all know what you’re talking about.

    But the big companies absolutely did not get permission to do this. Nor did anyone give permission to allow for the recent jailbreaks by AI task systems to hack into other companies. If that were any singular human (not of the owner class) they’d face charges under the CFAA. But since it’s a massive company with an army of blue-haired lawyers, nothing was done about it, and the companies actually bragged about their AIs escaping containment.

    Extinction by AI takeover is far more interesting than extinction by global drought.

  • busted_Anoose@aussie.zone
    link
    fedilink
    English
    arrow-up
    19
    ·
    7 days ago

    Yes but Aaron made the rookie mistake of not being a multibillionaire with a whole building of lawyers to fight on his behalf.

  • artyom@piefed.social
    link
    fedilink
    English
    arrow-up
    15
    ·
    9 days ago

    Aaron was not charged with scraping public servers, he was charged with unauthorized access. He “hacked” the server. Meta just scraped publicly accessible information.

    • DillDough@lemmy.zip
      link
      fedilink
      English
      arrow-up
      37
      ·
      9 days ago

      Bruh, can you lick metas boots any harder? There’s so many examples of them using unauthorized access even long before llm’s were a thing…but even if that wasn’t the case, they just fucking bragged about their “ai” hacking other companies. Stop spreading lies for the billionaires, it’s pathetic.

      • Rimu@piefed.social
        link
        fedilink
        English
        arrow-up
        8
        ·
        9 days ago

        Take it easy. artyom might just be mistaken you don’t need to turn it into a class war.

        • DeathsEmbrace@lemmy.world
          link
          fedilink
          English
          arrow-up
          20
          ·
          9 days ago

          At this point it already has been turned into a class war. They get away with so many crimes and just a slap on the wrist. If an individual does it life sentences. Almost like the laws are made to be broken by the rich.

          • Rimu@piefed.social
            link
            fedilink
            English
            arrow-up
            1
            ·
            8 days ago

            I’m not pretending there isn’t a class war. That would be silly of me.

            When people start acting like the class war is happening right here between their fellow fedi citizens, that’s when it becomes a bit unhinged.

            • raspberriesareyummy@lemmy.world
              link
              fedilink
              English
              arrow-up
              2
              ·
              7 days ago

              Between equals it’s not a class war though, it’s infighting. And agreed, we should avoid that. Rich parasites love to see us divided.

      • artyom@piefed.social
        link
        fedilink
        English
        arrow-up
        6
        ·
        edit-2
        9 days ago

        I am not “licking boots”, I am discussing facts. And I’m gonna keep doing it.

        There’s so many examples of them using unauthorized access even long before llm’s were a thing

        Sure, there’s an endless list of awful things they’ve done, but that’s not what’s being discussed here.

        • DillDough@lemmy.zip
          link
          fedilink
          English
          arrow-up
          7
          ·
          9 days ago

          I’m not talking about random awful things they’ve done I’m talking specifically about unauthorized access issues which is literally what you named as being the difference. Get some reading comprehension.

          • artyom@piefed.social
            link
            fedilink
            English
            arrow-up
            2
            ·
            edit-2
            8 days ago

            I’m talking specifically about unauthorized access issues which is literally what you named as being the difference.

            Correct, that is the difference between what I said and what the author said. The author is talking about scraping. Get some reading comprehension. You don’t even have to read the article, it’s literally in the title.

      • Blue_Morpho@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        ·
        9 days ago

        I’m on Aaron’s side, but didn’t he break into a network closet and patch directly into their servers?

        It’s splitting hairs but the law treats breaking and entering very differently than remote access. When I went to school, I had a job on campus but I’m sure the University would have considered it illegal if I broke into the utility closet of the building I worked in.

        • IsoKiero@sopuli.xyz
          link
          fedilink
          English
          arrow-up
          10
          ·
          9 days ago

          didn’t he break into a network closet and patch directly into their servers?

          No. He placed a computer in a unlocked network closet and used that to download a bunch of data from JSTOR (which he had acces to, and a lot of the data is public domain anyways) and kinda-sorta caused DDoS attack against the service. Prosecution then slapped him with a shitload of federal charges which (in my opinion) were largely at least massively exaggerated if not straight made up. Threat of 50 year jail sentence and million(s) in fines then pushed him to take his own life.

          Wikipedia has pretty detailed info about him and the whole case.

      • artyom@piefed.social
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        7 days ago

        If you want to be clear, you should provide some sort of explanation and evidence.

    • John Richard@lemmy.world
      link
      fedilink
      English
      arrow-up
      23
      ·
      9 days ago

      But he did have access via university. Besides, my understanding is he was scraping research journals, many of which would not have been possible without grants from tax dollars. Meaning the mere fact that the research is publicly funded means the journals should also be public.But they didn’t even prove he distributed them. For all we know, he was planning to have an AI use them for training and apparently that isn’t a problem.

      You do realize as well that almost every site you visit makes API requests. And you can then use the same API to request other data directly. And the difference between “public” and unauthorized access often comes down to whether you used the same API that your browser would call, but instead decided to make other calls to it directly.

      • JasonDJ@lemmy.zip
        link
        fedilink
        English
        arrow-up
        3
        ·
        edit-2
        9 days ago

        but instead decided to make other calls to it directly.

        Like with command line tools? But that’d be hacking! Egads!

        Is it a B&E if the door is wide open…or would that just be an ‘E’?

        If you let your neighbor walk in to your house as he pleases…can you get mad at him if he walks in like Cosmo Kramer while you’re shagging your wife? Assuming you’re not into that, of course.

        What if he had noble intents…like he heard the moans and thought you were at work and he was about to break up an affair or stop a rape or have some leftovers?

        • John Richard@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          8 days ago

          No, a much better comparison is that his backyard is already wide open. Whenever you go there he takes you all around his backyard. There are no signs saying you shouldn’t go to certain parts or fences stopping you. You go there to look at other parts and he still takes you there. Then one day he gets mad and says he himself took you to a part of his backyard that he didn’t want to, and rather than put up a better fence, he’s going to have you arrested for trespassing with multiple felonies… no warning, straight to prison!

  • DupaCycki@lemmy.world
    link
    fedilink
    English
    arrow-up
    15
    ·
    7 days ago

    When a regular person murders someone - life in prison.

    When We Love Genocide LLC murders 12 families - $3.43 fine.

    • Regna@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      8 days ago

      He (who founded and stood behind the original concept of Reddit) was so forcefully charged that he decided to commit suicide because he believed that information deserves to be free.

  • AlteredEgo@lemmy.ml
    link
    fedilink
    English
    arrow-up
    10
    ·
    9 days ago

    Obviously fuck Meta, but the difference is the publishing. Scraping and copying for personal or business use is a civil matter. And machine learning from unlicensed material is also fine - as long as the material isn’t “memorized” and an AI model can’t reproduce it.

    Aaron bravely published the papers which is a criminal matter. That is why I believe copyright and IP law is the real theft, they take down copies for people to learn from.

    There are many many millions of people who read and learned from pirated textbooks and who use those skills to do things. Who watched pirated amines or comics use that to learn how to draw.

    Basically we should not be siding with the unethical side of IP law just to oppose AI corporations. They can afford to buy the books and media, and a simple purchase will do. And they can afford the lawsuits. And they will figure out the memorization problem, so that AI models learn but not memorize (which is only happens in like 1% of the cases and only when you specifically ask for something “just like that”).

    China (so far) is saying that the AI models they produce should be open weight and be available to all people on earth. So if we have any issues it should be with the monopolization of AI models that concentrate this new developing immense power of AI in the hands of a few plutocrats. Which the IP laws might actually help them with.

    • rustydrd@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      8
      ·
      edit-2
      8 days ago

      He didn’t publish them, as far as I know. The charges brought against him were centered around the allegedly “fraudulent” use of his JSTOR account and the fact that he downloaded the files from the premises of an institution he did not belong to by plugging his laptop in a network switch in a restricted area where he wasn’t allowed to be. The actual deed was minor, likely not even criminal as far as digital rights were concerned, and the charges were famously so out of proportion that even other attorneys and legal scholars questioned them publicly. The publishers wanted to make an example, and it drove Swartz into suicide before a proper trial could be held.

      Meta and other companies do essentially the same thing at a much larger scale and with the clear intent to monetize it through publishing AI models built on top of it all. The main difference is the legal climate, which has changed since/due to Swartz, and the lack of clarity that the law has for these new applications. Substantively though, there is a good bit of hypocrisy here, and it makes sense to point this out.

    • technocrit@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      2
      ·
      8 days ago

      Basically we should not be siding with the unethical side of IP law just to oppose AI corporations.

      Why not? There is nothing ethical under capitalism. The system is literally destroying the planet. Normal people need to use whatever tools they can just to have a chance.

      • AlteredEgo@lemmy.ml
        link
        fedilink
        English
        arrow-up
        2
        ·
        8 days ago

        Lets say IP law is extended or reinterpreted to include that machine learning from book or papers or articles or comments requires a special license, even assuming the memorization problem is solved. This is what anti ai seems to be arguing.

        This would then result in some kind of “deal”. Producers of AI models are required to pay some kind of overall license fee or percentage into some kind of public fund. Even Bernie Sanders suggested something like that. The problems I see:

        1. The AI corporations can afford this and it will not really impact them in any way. The prices for AI rise a little. Free access is reduced.
        2. Open Weight models may no longer be used freely. You could still pirate a Chinese one but while capitalists have access to any potential benefits in replacing labor with AI, for ordinary people it becomes an additional form of rent.
          Also in combination with advances in robotics, these AI models could do a lot to allow people to become “independent” by just telling your $6000 robot (actual price for a humanoid robot in china today) to plant some potatoes and vegetables there there and there, then clean the house etc. Or DIY build your own out of plywood, servos and a smartphone once technology advances.
        3. The government through that fund gains an active interest in protecting that source of incoming and increase AI use, even if it does replace workers.
          This happened with the tobacco funds when vaping came around. Governments had leveraged the future payouts of these funds with banks for short term payouts and would have been in big trouble if vaping actually reduced smoking significantly.
      • kestrel7_7@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        8 days ago

        I thought that was where a bunch of the initial core of libgen/anna’s archive was from? Maybe I’m wrong