One of the reasons young people struggle so much with mental health nowadays is because there is a crisis of authority in society - no one actually tells them how to live well, with any suggestion of a normative or “right” way of doing things being deemed problematic

Instead they have nothing to guide them except deceptive inner monologues and colourful infographics made by San Francisco designers with no life experience telling them “you do you”, “it’s okay to be lazy”, “it’s okay to be selfish” on their instagram feeds
And so people end up making mistakes and suffering consequences that could have been prevented had they been given actual advice acquired over centuries of culture and tradition (Scruton’s “answers that have been discovered to enduring questions”)
But tradition alone is not enough - it needs to be embodied. We need teachers, counsellors, priests, etc. who are personally invested in our self-betterment and can tell us when we are doing things wrong from a place of love
Insisting that people will be happy if they just do whatever they feel like is not loving. It’s also a sentiment that has so obviously been exploited by capitalism since the 60s: a society that promotes unrestrained self-actualisation is the ideal consumer society

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like

My top 10 tweets of the year

A thread 👇

https://t.co/xj4js6shhy


https://t.co/b81zoW6u1d


https://t.co/1147it02zs


https://t.co/A7XCU5fC2m
First update to https://t.co/lDdqjtKTZL since the challenge ended – Medium links!! Go add your Medium profile now 👀📝 (thanks @diannamallen for the suggestion 😁)


Just added Telegram links to
https://t.co/lDdqjtKTZL too! Now you can provide a nice easy way for people to message you :)


Less than 1 hour since I started adding stuff to https://t.co/lDdqjtKTZL again, and profile pages are now responsive!!! 🥳 Check it out -> https://t.co/fVkEL4fu0L


Accounts page is now also responsive!! 📱✨


💪 I managed to make the whole site responsive in about an hour. On my roadmap I had it down as 4-5 hours!!! 🤘🤠🤘