SoundNativ

Machine Learning

Speech Machine Learning Engineer

  • Machine Learning
  • ๐ŸŒ Remote
  • Full-time
Apply for this role

When a learner says a word in SoundNativ, we tell them which sound was off and how to fix it, in about a second. That feedback is the product. If it is wrong, nothing else matters.

You will own the models and pipelines behind it: phoneme recognition, pronunciation scoring and accent analysis, from training run to production endpoint.

What you'll work on

  • Training and fine-tuning speech models (wav2vec 2.0, Whisper-style encoders and friends) for phoneme-level scoring and accent detection.
  • Deploying them behind our Python and FastAPI scoring service for real-time and batch inference.
  • Building evaluation sets that cover the first languages our learners speak, and tracking how scores hold up across them.
  • Data pipelines for audio and text: cleaning, labelling, augmentation and versioning.
  • Exploring where LLMs and speech-to-speech models can coach, not just score.

Who you are

  • Strong foundations in machine learning, statistics and signal processing, and pride in models that make it to production.
  • Equal parts creative and methodical. You try the proven and the unproven, and you measure both.
  • You care about the person on the other end. A score that feels unfair to a Korean speaker is a bug.
  • You like long stretches of focused work without meetings, and you don't wait for permission.

Requirements

  • 3+ years training, evaluating and deploying ML models in production, ideally in speech or audio.
  • Fluent in Python and PyTorch.
  • Experience with ASR, forced alignment or phoneme recognition.
  • Comfortable owning infrastructure: GPUs, inference latency and cost.

Nice to have

  • Background in phonetics, or a good ear for accents.
  • Published research or open-source work in speech.
  • Experience optimising models for low-latency or on-device inference.

What we offer

  • Competitive pay, based on your experience.
  • Work from anywhere.
  • Flexible hours. We care about what you ship, not when you are online.
  • Real ownership from week one, and a direct line to the founders.
  • The AI tools you need to move fast, paid for.

About SoundNativ

SoundNativ is an accent training app for people who speak English as a second language. Learners record themselves, get sound-by-sound feedback from our speech models in about a second, and practise with lessons built around their first language. More than 70,000 people have used it to sound clearer at work, in interviews and in everyday life, and they rate it 4.8 stars.

How we work and how we hire โ†’

Apply

Apply for Speech Machine Learning Engineer

About five minutes. No cover letter needed, just your CV and one honest answer.

City and country, so we can plan around time zones.

GitHub, Hugging Face, Google Scholar or a write-up of your work.

PDF, up to 5 MB.

0 / 4,000

Other open roles