Discussion about this post

User's avatar
Mai Frances's avatar

As ever I’m late to the party, just discovered the work of Riddance and this substack! This post piqued my interest, as an academic in astrophysics I know absolutely nothing about social media algorithms but I do know a bunch about scientific coding (read: throwing stuff at a wall and seeing what sticks) and the burden of statistical significance and this seems like the kind of data science that is prolific across the field of astrophysics (and probably many other academic fields, much to the horror of actual data scientists): how do you classify and infer latent hierarchy in a large, heterogenous dataset without knowing enough about the underlying generative process to model the likelihood function? The answer always seems to be to fucking about with clustering algorithms and seeing if you stumble upon any cool hierarchies in your galaxies (or whatever). But seems like the same principle could be applied here.

I abandoned Instagram a while ago, prior to AI generated video, so this may be a naive comment but do you have a feel for the proportion of authentic vs inauthentic accounts in these niches, is it possible there’s bias being introduced due to a skewed sample? And have you considered using an API scraper to pull all manner of account information for entire searches or hashtags (from an admittedly quick and dirty google search, it seems like there is a tool that reliably does this, although with a cost)? You could probably do a whole lot of probabilistic analysis on various metadata to start building a picture of whether there are any underlying patterns through which this kind of content could be classified for prioritisation in the algorithm fairly simply using python (although astronomers assume everything is Gaussian so i am no authority on how astronomy stats measure up in real world applications!)

No posts

Ready for more?