Academics routinely use the term “sub-Saharan Africa”. Until recently, I also used it without really thinking much about it or questioning its use. But it’s problematic and should be avoided. Here are two excellent pieces explaining why:

It’s confusing and geographically inaccurate:
https://t.co/j3qZgTy3IT
"46 of Africa’s 54 countries are listed as “sub-Saharan,” excluding Algeria, Djibouti, Egypt, Libya, Morocco, Somalia, Sudan & Tunisia. This doesn’t make geographical sense: 4 countries included are on the Sahara, while Eritrea is “sub-Saharan”, but its neighbour Djibouti isn’t.”
It’s rooted in racist thought:
https://t.co/IsVIk81D3C
“it divides Africa according to white ideas of race making North Africans white enough to be considered for their glories, but not really white enough…It is a way of saying “Black Africa” and talk about black Africans without sounding overtly racist.”
Instead use more accurate geographic markers like East, West, Central and Southern Africa or simply list the names of the countries you are refereeing to.

More from All

How can we use language supervision to learn better visual representations for robotics?

Introducing Voltron: Language-Driven Representation Learning for Robotics!

Paper: https://t.co/gIsRPtSjKz
Models: https://t.co/NOB3cpATYG
Evaluation: https://t.co/aOzQu95J8z

🧵👇(1 / 12)


Videos of humans performing everyday tasks (Something-Something-v2, Ego4D) offer a rich and diverse resource for learning representations for robotic manipulation.

Yet, an underused part of these datasets are the rich, natural language annotations accompanying each video. (2/12)

The Voltron framework offers a simple way to use language supervision to shape representation learning, building off of prior work in representations for robotics like MVP (
https://t.co/Pb0mk9hb4i) and R3M (https://t.co/o2Fkc3fP0e).

The secret is *balance* (3/12)

Starting with a masked autoencoder over frames from these video clips, make a choice:

1) Condition on language and improve our ability to reconstruct the scene.

2) Generate language given the visual representation and improve our ability to describe what's happening. (4/12)

By trading off *conditioning* and *generation* we show that we can learn 1) better representations than prior methods, and 2) explicitly shape the balance of low and high-level features captured.

Why is the ability to shape this balance important? (5/12)

You May Also Like