MLOps Engineer

Fathom.io


Date: 8 hours ago
City: Dhahran
Contract type: Full time
About The Role

We’re looking for a mid-to-senior MLOps engineer to help build the intelligence layer of our AI platform. This is not a traditional MLOps role focused only on maintaining pipelines and deployments. You will help create the underlying infrastructure that makes sophisticated AI workflows accessible through a simple, self-service product experience: model deployment, training, notebooks, functions, UI-driven agent creation, and RAG pipelines.

You’ll work at the intersection of platform engineering, machine learning infrastructure, distributed systems, and developer experience.

What You’ll Do

  • Design and build the infrastructure powering our Intelligence layer.
  • Enable reliable, automated workflows for model training, deployment, lifecycle management, and inference.
  • Build scalable foundations for users to create, configure, and operate AI agents and RAG pipelines through the platform UI.
  • Develop the platform capabilities behind managed notebooks, functions, experiments, training jobs, model registries, and serving endpoints.
  • Improve model serving, observability, versioning, evaluation, promotion, and rollback capabilities.
  • Optimize GPU inference and training deployments for performance, reliability, and cost efficiency.
  • Explore efficient approaches for deploying models across centralized GPU infrastructure and edge devices.
  • Automate workflows that otherwise require manual MLOps effort, with a focus on safe, self-service capabilities for platform users.
  • Partner with backend, product, and AI teams to turn complex infrastructure into intuitive platform features.
  • Help define standards for security, multi-tenancy, resource isolation, model governance, and operational reliability.

Our current stack

You’ll Work With Technologies Including

  • Kubernetes and Knative
  • KServe and vLLM
  • MLflow
  • LangFuse
  • GPU infrastructure and model-serving workloads
  • RAG architectures, vector retrieval, agent workflows, and LLM applications

Our intelligence backend is built in Rust, so Rust experience is a strong advantage.

What We’re Looking For

  • Strong experience in MLOps, AI platform engineering, machine learning infrastructure, or distributed systems.
  • Hands-on Kubernetes experience, including deploying and operating stateful, training, notebook, serverless, or GPU-intensive workloads.
  • Experience building ML platforms that support some combination of training, experimentation, notebooks, feature/data workflows, model registries, and production serving.
  • Experience with model serving frameworks such as KServe, vLLM, Triton, Ray Serve, or similar.
  • Familiarity with ML lifecycle tooling such as MLflow, experiment tracking, model registries, and CI/CD or GitOps for ML systems.
  • Practical experience optimizing training and inference workloads for latency, throughput, availability, and cost.
  • Understanding of GPU scheduling, resource allocation, model quantization, batching, autoscaling, and multi-model serving.
  • Experience with RAG systems, LLM applications, AI agents, or their supporting infrastructure.
  • Strong software engineering fundamentals and a product-minded approach to platform design.
  • Ability to make complex operational workflows reliable and approachable for users.

Nice to have

  • Rust experience or interest in working with Rust-based backend systems.
  • Experience running notebook environments such as JupyterHub, Kubeflow Notebooks, VS Code Server, or equivalent.
  • Experience creating secure, scalable execution environments for user-defined functions or jobs.
  • Experience deploying AI workloads on edge devices.
  • Familiarity with model optimization techniques, including quantization, compilation, pruning, and hardware-aware serving.
  • Experience with multi-tenant AI platforms, security controls, and model governance.
  • Familiarity with observability and evaluation tooling for ML and LLM systems.

What Success Looks Like

You will help us evolve from manually operated AI infrastructure to a platform where users can train, evaluate, deploy, and operate models; work in managed notebooks; run functions; create agents; and configure RAG workflows - with the platform handling as much of the operational complexity as possible.

By submitting this application, I agree that my personal data will be collected, processed, and retained by the company solely for the purposes of managing and assessing my candidacy.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.

Post a resume

Similar jobs

Field HSE Specialist

SLB, Dhahran
1 day ago
The Field HSE Specialist is responsible for supporting local managers in establishing and continuously improving the quality and health, safety and environment (HSE) culture at the worksite.Develop awareness and ensure quality and HSE compliance are an integral part of Line Management responsibilities and objectives.Assist management and support functions in: implementing and improving the Quality and HSE Management Systems, and defining...

ASSOCIATE CYBERSECURITY ANALYST

Johns Hopkins Aramco Healthcare (JHAH), Dhahran
2 days ago
Job Description SummaryThe Associate Cybersecurity Analyst is responsible for implementing and maintaining the organization’s cybersecurity framework, utilizing their working knowledge in at least one of the relevant technologies and processes to protect sensitive data and guard against any malicious disruption of business continuity at JHAH.Minimum Education RequiredBachelor's Degree in the field of Information Technology or in a relevant field.Professional Certifications...

Skilled Labors (Non-National)

Front End Limited Company, Dhahran
4 weeks ago
Role Purpose The Skilled Labor (Non-National) provides essential hands-on support to maintenance and operations crews — material handling, site preparation, assistance to Technicians, and housekeeping — executed safely and efficiently in line with site HSE requirements.Key ResponsibilitiesAssist Technicians in maintenance tasks: holding, positioning, passing tools, and supporting lifts.Handle, move, and stage materials, spare parts, and equipment at work sites and...