Job Description
Edison Smart is supporting a fast-growing global technology business that is continuing to invest heavily in its AI and machine learning capabilities.
We are looking for a Senior LLMOps Engineer to take ownership of key infrastructure across LLM training, inference and data management.
This is a hands-on engineering position suited to someone who combines strong software engineering with MLOps/platform expertise and has experience building infrastructure that supports production AI systems.
The Role
You will work closely with AI/ML and engineering teams to build and improve the platforms used to train, deploy and operate large language models at scale.
Responsibilities will include:
- Build and improve infrastructure for LLM training, fine-tuning and inference.
- Develop platforms for training job management, resource scheduling and model lifecycle management.
- Deploy and operate open-source LLMs alongside integrations with external model APIs.
- Build centralised capabilities around model access, permissions, quotas and request management.
- Develop training data infrastructure covering storage, versioning, quality, access control and traceability.
- Improve GPU utilisation and optimise the performance, reliability and cost of AI workloads.
- Build automation around model deployment, monitoring and infrastructure operations.
- Troubleshoot complex production and infrastructure issues.
- Participate in an on-call rotation supporting production AI services.
- Work closely with AI/ML, engineering and security teams to support wider adoption of AI across the organisation.
What We're Looking For
- 5+ years across Software Engineering, Platform Engineering, SRE, MLOps or AI Infrastructure.
- Strong Python engineering skills.
- Hands-on experience with model training, fine-tuning and/or inference infrastructure.
- Experience deploying open-source LLMs and integrating third-party LLM APIs.
- Strong Linux and Kubernetes knowledge.
- Experience working with GPU infrastructure or compute-intensive workloads.
- CI/CD and observability experience.
- Understanding of model and dataset storage, versioning and access control.
- Strong understanding of production engineering principles including scalability, reliability and performance.
- Comfortable operating as a senior individual contributor while providing technical leadership across teams.
- Fluent English.
Nice to Have
- Distributed model training experience.
- LLM inference optimisation.
- GPU cluster management.
- Contributions to open-source AI infrastructure or MLOps projects.
- Experience operating AI/ML platforms at significant scale.
- Chinese language capability.