Back to Jobs
E
Deadline Passed

Senior ML Engineer - Model Training

eMumbaPK
Last Date to Apply
13 September 2026
Date Posted
14 August 2026
Total Views
0 Candidates

Job Description

We are looking for a hands-on Machine Learning Engineer to own the training, fine-tuning, and continual improvement of the models that power our metrics. You will build and curate datasets, optimise encoder and decoder models, and run systematic experiments across architectures, batch sizes, and training regimes to improve accuracy, efficiency, and cost-performance trade-offs.

You will also own the end-to-end deployment pipeline, packaging trained models into production-ready images running on NVIDIA’s inference stack and maintaining the Rust services that serve them in production. This is a highly technical role with direct ownership from dataset construction and experimentation through to production deployment.

Key Responsibilities

• Model training and fine-tuning (top priority): Train, fine-tune, and improve encoder

and decoder models for our metrics. Own the full loop from data to evaluated,

production-ready checkpoints.

• Dataset and metric development: Design, build, and curate datasets for new and

existing metrics. Define labelling schemes, manage data quality, and connect dataset

changes to measurable model improvements.

• Experimentation and evaluation: Run systematic experiments across model

architectures, batch sizes, precision, and training hyperparameters. Build and maintain

rigorous evaluation harnesses and track results to find the best accuracy/cost trade-offs.

• Inference optimisation: Optimise trained models for production inference on the Nvidia

stack (Triton, TensorRT / TensorRT-LLM, ONNX) using methods such as quantization,

distillation, and precision tuning.

• Model deployment pipeline: Maintain and improve the pipeline (Docker-based) that

converts models to suitable formats (TensorRT, ONNX), fixes vulnerabilities, and

containerises them for production.

• Rust application layer (supporting): Maintain and extend the Rust services and API

endpoints that serve our models in production.

Skills, Knowledge and Expertise

Must-Have Skills

• 3+ years of practical experience training and fine-tuning deep learning models, with a proven

track record of taking models from data to production

• PyTorch, and a solid understanding of the inner workings of Transformers and deep

learning more broadly

• Dataset construction and model evaluation: building datasets, defining metrics, and

designing evaluation methodology

• Hands-on Python programming experience

• Nvidia Inference Stack: Triton Inference Server, TensorRT / TensorRT-LLM, ONNX

• Docker and Kubernetes (k9s familiarity is a plus)

Nice-to-Have

• GPU model training and experiment tracking (e.g. Weights & Biases, MLflow)

• Inference optimisation techniques (e.g. quantization)

• Backend development in Rust, including building and maintaining production API

services

Soft Skills

• Builder mindset: thrives on writing, debugging, and improving production code.

• Collaborative, humble, and open to feedback.

• Strong communicator who explains design decisions clearly.

• Influences through contribution, not hierarchy.

About Emumba

Emumba is a global engineering and consulting company with strengths in software development and an established AWS cloud practice focused on Data and GenAI. For 15 years, our teams across the US, the UAE, and Pakistan have earned trust through quality delivery and ownership of work. We look for people who value the culture they work in as much as the craft they bring to it.

More Options

JobsConsultingLibraryLogin