How SL2T Actually Works: Landmarks, Not Gloss, Not Raw Video
SL2T does not see sign language the way a video call does. An on-device model first converts a signer's hand, face, and body movement into geometric coordinate landmarks - MediaPipe Holistic tracks 130 key points across the face, body, and hands, frame by frame - and only those coordinates, never the original camera feed, are sent off for translation [1][2]. The camera feed itself is discarded immediately, a design Google frames explicitly as a privacy safeguard rather than an afterthought [1].
The more consequential design choice sits one layer deeper: SL2T translates directly from that landmark sequence to text, skipping the intermediate 'gloss' annotation step that most prior sign-language translation systems relied on [3]. Glosses - a word-for-word transcription convention - fail to capture rich, non-linear aspects of sign languages such as non-manual markers and spatial constructions [3]. Google also says the model is specifically optimized for left-handed and one-handed signing, noting roughly 10% of signers are left-handed [4], and includes hallucination-prevention mechanisms designed to avoid falsely transcribing non-signing movements as sign input [5]. These are the kinds of details that rarely make a press release but are exactly what determines whether a translation model is usable in daily life rather than just a lab demo.


![[8/12 15:00] Made by Google 2026 - Pixel 11, Gemini, and the SL2T sign language model / CodeRabbit raises $143M](https://img.youtube.com/vi/lB5Gej3Yf4c/mqdefault.jpg)
