My favorite example of how informationally toxic YouTube's algorithm is this:

Imagine you're high school freshman and got a school assignment about the Federal Reserve.

You watch videos on YouTube all the time, so you go home and put "Federal Reserve" into YouTube's search bar.

This is the first video that comes up (1.6 million views) https://t.co/vt6GQugkHa
It is, perhaps not surprisingly, conspiratorial quackery.

Once you view that, the algorithm also then suggests this on the sidebar:

"What You're Not Supposed to Know About America's Founding" (863,00 views)

https://t.co/anOLNuPJGF
It's a John Birch Society lecture delivered by someone who says he "infiltrated Marxist organizations as a young man" and that Marxists infiltrate the media.
He then delivers a convoluted lecture about the Illuminati being behind both both global communism and the abolitionists. Charles Sumner and Horace Greeley were communists, and the Civil War was a communist/Illuminati plot, started by communist terrorist John Brown.
Once you've viewed that, you can then click on this video, conveniently furnished by the algorithm:

Trump Tells Everyone Exactly Who Created Illuminati (4 million views)
https://t.co/2ipMBtaphA
And on and on and on.

All this starts with someone genuinely searching for basic information about the Federal Reserve!

And this is where you can end up in two or three clicks.

And it's making money for YouTube all along the way.

Something's deeply rotten here.

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like