
Senior Computer Vision Engineer at Sword Health building AI to understand human movement in real time.
Thrive is Sword's program for chronic joint and back pain: personalized, clinician-designed physical therapy delivered at home, with an AI care specialist guiding members between visits. What makes it work is movement understanding. During each session, Thrive reads how a member moves from the camera and gives real-time feedback on their exercises, the way a physical therapist in the room would.
The Computer Vision team, part of the Algorithms org, builds the models behind that. We turn a camera feed into an accurate read of human movement, 2D and 3D pose and body dynamics, in real time, on-device or in-the-cloud, and turn that movement into clinical signals. This role owns that computer vision end to end: the models, the data lifecycle that feeds them, the evaluation, and the systems that ship them to members at scale.
Own core computer vision models, from 3D human pose to statistical body modeling, taking them from prototype to production;
Ship those models to run real-time in the cloud and on-device on tablets, owning the conversion and optimization in between;
Own the data and code lifecycle behind them: training frameworks, annotation workflows, pipelines, auto-labeling, test sets and taxonomy;
Extend movement understanding into multimodal territory, combining it with language and reasoning; build novel approaches in the movement-intelligence domain;
Unify and mature how we train, track, version and deploy models, so every result is reproducible and testable;
Design systems that run without you, automating the loops so the team's output scales past manual effort;
Help grow the Computer Vision team by defining and promoting best practices, establishing principles that scale your impact.
5+ years solving complex problems with Computer Vision, with models shipped to production;
Strong software engineering foundation across architecture, pipelines, MLOps and the full model lifecycle;
Deep, hands-on command of modern Computer Vision (transformers and convolutional models), with real depth in at least one of detection, segmentation, tracking, pose, or 3D;
A data-centric instinct: you cook your own data, build data flywheels and active-learning loops, and treat the dataset as source code;
Solid grounding in multimodal AI, with the ability to build systems that pair vision with language and reasoning when the problem calls for it;
Strong written, asynchronous communication. You think in specs and documents and leave a clear trail others can build on;
Self-direction and full ownership: you scope ambiguity into a plan and ship without being told;
Fluency with PyTorch or JAX, the Python data stack, and comfort picking up new languages and codebases;
AI-native ways of working: you build AI into the workflow itself, from agents to LLM-in-the-loop tooling.
Experience with pose estimation, body modeling, movement intelligence, or human-motion work;
Experience working with massive data, building full-lifecycle ML systems at a startup or scaleup, wearing different hats;
Experience automating experimentation or model-improvement loops;
Product mindset, users empathy and desire to ship impact to millions of members.