
Job Summary
We are seeking a Senior AI/ML Engineer to design, deploy, operate, and optimize production-grade AI and Machine Learning solutions at scale. This role focuses on LLM/Generative AI applications, MLOps, cloud platforms, and software engineering excellence, ensuring AI systems are reliable, observable, cost-efficient, secure, and production-ready. You will work closely with data scientists, platform engineers, and product teams to build and support RAG, agentic AI, and ML solutions across the full development lifecycle.
Key Responsibilities
- Design, build, deploy, and support production AI/ML applications, including LLM-powered, RAG, and agent-based solutions.
- Develop robust evaluation, testing, observability, and monitoring frameworks for AI systems.
- Implement and maintain CI/CD pipelines for ML and GenAI workloads.
- Monitor and optimize model performance, latency, reliability, cost, and operational health.
- Build and manage AI infrastructure using Infrastructure-as-Code and cloud-native services.
- Troubleshoot production issues across models, data pipelines, retrieval systems, agents, and integrations.
- Collaborate with engineering, data science, and platform teams to deliver scalable AI solutions.
- Drive engineering best practices including code reviews, testing, version control, and documentation.
- Implement governance, guardrails, tracing, logging, and monitoring to ensure responsible AI deployment.
- Mentor junior engineers and contribute to technical leadership within the team.
General Qualifications
- Bachelor's degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or a related discipline.
- 5+ years of experience in Software Engineering, Machine Learning Engineering, MLOps, or AI Engineering roles.
- Experience designing and supporting production systems in cloud environments.
- Strong communication, stakeholder management, problem-solving, and mentoring capabilities.
Mandatory Requirements
- Strong Python programming expertise with experience building and maintaining production-grade applications.
- Solid software engineering fundamentals, including testing, code reviews, Git/version control, and maintainable code practices.
- Proven experience delivering and supporting LLM/Generative AI applications in production.
- Hands-on experience with RAG architectures, AI agents, and/or fine-tuned LLMs.
- Strong understanding of LLM evaluation, guardrails, observability, latency optimization, and cost management.
- Experience implementing Infrastructure-as-Code using Terraform or equivalent IaC tools.
- Production experience on at least one major cloud platform (Azure, AWS, or GCP).
- Experience with Databricks or comparable lakehouse/MLOps platforms.
- Hands-on experience with Docker and Kubernetes for containerized AI/ML workloads.
- Experience building and supporting CI/CD pipelines for ML and AI deployments.
- Strong knowledge of LLM tracing, logging, telemetry, and observability frameworks.
- Experience implementing monitoring solutions using tools such as Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, or Azure Monitor.
Nice-to-Have Skills
- Experience with TensorRT-LLM, FlashAttention, or other LLM inference optimization technologies.
- Knowledge of tensor parallelism and pipeline parallelism for large-scale model deployment.
- Experience with AI orchestration frameworks such as LangGraph, LlamaIndex, AutoGen, or Semantic Kernel.
- Familiarity with Model Context Protocol (MCP).
- Experience with LLMOps tooling, including LiteLLM, model routing/fallback strategies, prompt/version management, and token cost monitoring.
- Experience with workflow orchestration platforms such as Airflow, Dagster, Kubeflow, or Argo.
- Relevant cloud, AI, Kubernetes, or Databricks certifications.
- Previous experience in consulting, professional services, or client-facing delivery environments.