All open positions

Principal ML Systems Engineer

On-siteAustin, TexasFull-time

About Throne

Throne is a continuous tracking device for getting personalized insight about gut health and hydration. Co-founded by John Capodilupo — co-founder and former CTO of WHOOP — we're bringing the rigor of continuous health tracking to two signals that have long been a guessing game. Our north star is to improve health and save lives.

Principal Engineer

We’re looking for a Principal Machine Learning Systems Engineer to own technically challenging problems spanning machine learning, data systems, backend infrastructure, and production software.

This is a hands-on role for an experienced engineer who thrives when the problem is important and the path to solving it is unclear. You will be expected to understand the objective, investigate the system and data, identify the highest-leverage work, make sound architectural and modeling decisions, and drive solutions from initial investigation through production deployment and operation.

At Throne, the boundary between ML and software engineering is intentionally porous. A model-performance problem may turn out to be a labeling problem, a data-pipeline problem, an inference architecture problem, or a product instrumentation problem. We want someone who is comfortable following the problem wherever it leads.

Location

Austin, Texas (In-person)

What You’ll Own

  • End-to-end technical problem solving. Tackle ambiguous product and engineering problems by identifying the critical questions, determining what needs to be built or learned, and driving the work to a measurable outcome.
  • Production ML systems. Build, train, evaluate, optimize, deploy, and monitor models across computer vision, video, audio, and other sensing modalities.
  • Backend architecture. Design and operate production services and APIs supporting device data, customer-facing applications, ML inference, and internal systems.
  • Large-scale media and ML pipelines. Own data ingest, object storage, queueing, worker orchestration, indexing, inference, dataset generation, and downstream metric computation.
  • Dataset and evaluation systems. Develop strategies for data mining, labeling, dataset quality, split design, active learning, model evaluation, regression testing, and failure analysis.
  • Model performance and efficiency. Drive improvements through model architecture, training strategy, transfer learning, quantization, distillation, inference optimization, and cost/performance tradeoffs.
  • Cloud infrastructure and deployments. Build and operate scalable infrastructure across compute, storage, queues, databases, networking, secrets, and model-serving systems.
  • Data architecture. Design schemas and data flows that support both product workloads and high-volume ML/data-science workflows.
  • Production reliability. Establish observability, metrics, alerting, capacity planning, incident response, rollout safeguards, and model-performance monitoring.
  • Sensor and algorithm experimentation. Work with hardware, firmware, product, and R&D teams to evaluate new sensing modalities and determine whether they create meaningful improvements in system performance.
  • Technical direction. Identify architectural weaknesses, model-performance bottlenecks, scalability constraints, and opportunities that may not already exist on the roadmap—and take responsibility for addressing them.

What We’re Looking For

  • 7+ years of professional engineering experience spanning software engineering, ML engineering, data infrastructure, or closely related disciplines.
  • A track record of owning complex technical problems from loosely defined objectives through production outcome.
  • Strong Python skills and substantial experience with production ML frameworks such as PyTorch.
  • Strong backend engineering experience building distributed production systems. Experience with Go is highly desirable.
  • Experience personally training and improving ML models, not simply integrating third-party model APIs.
  • Strong understanding of ML evaluation, dataset construction, failure analysis, experimentation, and model-performance tradeoffs.
  • Experience building large-scale asynchronous or data-processing systems using object storage, queues, worker fleets, and databases.
  • Strong SQL and relational-database fundamentals; PostgreSQL experience preferred.
  • Strong cloud infrastructure experience, ideally with AWS and infrastructure-as-code such as Terraform.
  • Experience deploying and operating ML inference workloads in production.
  • Strong software architecture and API-design instincts.
  • Ability to reason scientifically about ambiguous technical problems: form hypotheses, design experiments, interrogate data, and make evidence-backed decisions.
  • Exceptional technical independence. You know when to make a decision, when to gather more information, when to ask for input, and when to challenge the premise of the problem itself.
  • Strong communication skills and the ability to work effectively across backend, ML, firmware, hardware, product, and R&D.
  • Comfortable using modern AI/LLM development tools to increase the speed and quality of engineering, research, and analysis.

Nice to Have

  • Computer vision experience across classification, segmentation, detection, embeddings, and video/temporal models.
  • Audio or signal-processing experience.
  • Multimodal or sensor-fusion experience.
  • Model optimization using ONNX, TensorRT, quantization, distillation, or similar techniques.
  • SageMaker or other ML infrastructure experience.
  • Experience with labeling systems, active learning, weak supervision, or model-assisted annotation.
  • Experience with high-volume media processing.
  • Experience working with connected devices or embedded systems.
  • Experience validating new hardware sensors.
  • Experience operating systems with significant cost-per-inference or cloud-cost constraints.
  • Experience inheriting complicated systems and making them simpler and more reliable.

Why This Role Is Different

Most companies divide this work amongst backend engineers, ML engineers, data engineers, ML platform engineers, and research scientists.

At Throne, many of our most important technical problems cut directly across those boundaries.

We want someone who can start with an outcome: improve this model, make this pipeline scale, determine whether this sensor helps, reduce inference cost, understand why these sessions fail, and take responsibility for figuring out the rest.

You won't be expected to know the answer immediately. You will be expected to know how to find the answer, turn it into a plan, and execute that plan from start to finish.

Employment
Full-time
Work arrangement
On-site
Location
Austin, Texas