2. Of these, 34.6M parallel sentences were newly mined as a part of this work (~33M from IndicCorp that we released last year)
AI4Bharat/@iitmadras and EkStep Foundation announced the release of Samanantar, the largest publicly available collection of parallel corpora for Indic languages. This work was supported by Tarento Technologies and the @rbc_dsai_iitm Download:
2. Of these, 34.6M parallel sentences were newly mined as a part of this work (~33M from IndicCorp that we released last year)
4. A new single script joint model for En-Indic and Indic-En translation which outperforms all commercial and publicly available systems.
This is a tremendous effort that will have a great impact on Indic NLP research! @OfficialIndiaAI