Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)G
Posts
93
Comments
1020
Joined
3 yr. ago

  • I'm not surprised. I am surprised that the researchers were surprised, though.

    Bridging algorithms seem promising.

    The results were far from encouraging. Only some interventions showed modest improvements. None were able to fully disrupt the fundamental mechanisms producing the dysfunctional effects. In fact, some interventions actually made the problems worse. For example, chronological ordering had the strongest effect on reducing attention inequality, but there was a tradeoff: It also intensified the amplification of extreme content. Bridging algorithms significantly weakened the link between partisanship and engagement and modestly improved viewpoint diversity, but it also increased attention inequality. Boosting viewpoint diversity had no significant impact at all.

  • And what do I care about Reddit getting paid?

    If the IA doesn't complain about being used, then it's fine for me. The ideal outcome would be, if the archive can make some arrangement where they scrape the data and provide it to everyone. That way, sites only get scraped once and not constantly hammered.

  • Why?

  • But I’m talking about archived webpages and information previously available to the public with zero commercial value that has been removed.

    It is still "intellectual property". Maybe the policy is to just oblige removal requests if the content doesn't seem to be of public interest. Cause why not, right? Look at all the people here on Lemmy angry that their worthless posts are scraped or deleting them on Reddit. Obliging takedown requests is certainly the path of least resistance.

  • Ahh. Lovecraftian horror.

  • I don't know... I mean, I agree. But I'm seeing a lot of demands that instances should prevent scraping. Ok, it could be astroturf; a campaign by Reddit/data brokers to neutralize the free competition. But you have seen all those deleted posts on Reddit. Those are some special little minds.

  • Hmm. There are many things that could cause legal trouble for the Wayback Machine. I wouldn't jump to conclusions.

    You can see on Lemmy that many people would prefer to outlaw scraping, fair use, and all that. Well, not for the "good guys" obviously, but the law doesn't work on vibes. The IA would be legally impossible in most countries. In the EU, it would be a major crime because of copyright and GDPR. It's only the traditional US commitment to free speech and fair use that makes it possible at all.

    The IA exists in a legally precarious position. That's not because of any shady backroom dealing. If the crowd in this community had its way, it would be gone.

  • Other sites, eg with books and journals, are doing the same thing. They hope that they can extract more money by reducing the availability of their content.

  • I don't know their take-down policy. Could be privacy, could be copyright.

    I think they are shielded by Section 230 under US law. That means, if they don't do take-downs when requested, they become liable just like the original uploader. So it depends on whether they think they can defend something as fair use. IDK what they do with requests under non-US laws.

  • I doubt it's an honest mistake or simple hypocrisy. You can see that AI is both supposed to be useless and see hugely increased usage. Sure, people can be pretty dumb but this is really heavy.

    Well, whatever the reason for this may be... You will certainly not reason these accounts out of posting this stuff with numbers.

  • Saying it's not bad is too strong. All human activity has undesirable side effects.

    But yes. People who peddle that environment narrative are definitely not interested in improving matters.

  • Reddit can be scraped just as much as online books and journals.

  • Reddit is archived and available as torrent up until the API change.

  • Technologically no. Reddit sends out the data to 10s of millions of users as part of their normal operations. They need to try to block those who collect that data for the IA. Reddit has the very short end of the stick.

    The problem is that evading such counter-measures may be criminal in the US. Obviously, EU laws are much harsher.

  • In Icelandic ð cannot be used at the start of a word

    Didn't know that. I think it was fine in Old English.

    Yeah, phonetically they are different. I think they are using them correctly.

  • ðe ... þinking

    You are distinguishing eth and thorn and using them correctly? I am impressed; also a bit weirded out, but really impressed.

  • Lets say the beak-even cost for a single request is somewhere between $1-$5 depending on the request just for the electricity,

    Are you baiting the fine people here?

  • Yes. It is a big problem for Europe. I don't expect that it will be fixed in the foreseeable future. In fact, it is being made worse in many ways.

    You may reference and quote journal articles. That's something I expect will stay allowed.