Saudi scapegoats general for assassination of #Khashoggi:

By royal decree, Ahmed Assiri has been removed from his post as a deputy intelligence chief.

This is insane.

The Saudis brought a BONE SAW with them.

https://t.co/CtrnLR5lCB
“He got in a fight with 18 of our guys and accidentally died so then we totes had to cut up his body so we could transport his remains more easily.”

(KSA has not addressed the whereabouts of #Khashoggi in revealing the results of its “investigation.”)
https://t.co/r83J64P1Ln
Crown Prince MbS is using this to consolidate power.

(There were hopeful rumors the King would rein him in.)

This feels very Kushner-like, tbh.

https://t.co/GSVcIEKHNU
This thread is really what they’re going with. All in.

I anticipate the WH will accept this version that offers (barely) plausible deniability to MbS and celebrates his and the Kingdom’s “transparency” and “swift action,” or something to that effect.

https://t.co/WRJdCUbscz

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like