ENVIRONMENT:
Our client is seeking highly specialised AI Platform Engineers to design, build, operate and optimise enterprise-grade AI infrastructure within a complex, regulated environment. This is not a general cloud engineering, IT infrastructure or data science role. The successful candidates must have hands-on experience supporting production AI workloads across multi-cloud environments and must be comfortable working across AI infrastructure, model serving, agentic AI, security, observability, infrastructure-as-code and AI cost governance.
DUTIES:
· Design, deploy and optimise scalable multi-cloud AI platform infrastructure.
· Build reusable platform components for AI gateways, model serving, vector databases, data pipelines and GPU workloads.
· Develop and maintain infrastructure-as-code using Terraform, Pulumi, CloudFormation or equivalent technologies.
· Design infrastructure supporting agentic AI, including orchestration environments, tool-calling, agent memory, state management and multi-agent communication.
· Implement cloud-agnostic model-serving patterns that support workload portability.
· Define and manage AI platform SLAs covering availability, inference latency, throughput and reliability.
· Implement platform observability, monitoring, incident management, release management and operational runbooks.
· Design and implement zero-trust security controls for AI platforms.
· Manage AI compute expenditure through cost attribution, chargeback/showback, workload optimisation and usage reporting.
· Maintain technical documentation, architectural decision records and governance evidence.
· Mentor engineers and contribute to platform engineering standards and delivery practices.
REQUIREMENTS:
• Senior level: approximately 5–8 years of relevant cloud and AI platform engineering experience.
• Lead/Principal level: approximately 8–12 years of relevant experience, including technical leadership and responsibility for engineering teams or platform squads.
Mandatory Technical Experience:
• Candidates must demonstrate meaningful production experience in most of the following:
• At least two of the following AI ecosystems:
o AWS Bedrock or SageMaker
o Microsoft Azure AI Foundry or Azure OpenAI
o Databricks AI
o Enterprise Hugging Face deployments
• Kubernetes, Docker, Helm and containerised platform services.
• Terraform, Pulumi, AWS CDK, CloudFormation or equivalent infrastructure-as-code.
• CI/CD and automated deployment of cloud or AI platform components.
• Production model-serving infrastructure, AI gateways or inference endpoints.
• Platform observability using tools such as Prometheus, Grafana, Datadog, OpenTelemetry or Databricks Lakehouse Monitoring.
• Cloud security, identity and access management, including OAuth/OIDC, JWT, RBAC or ABAC.
• Production incident management, SLAs, release management and operational readiness.
• Experience within banking, financial services or another highly regulated enterprise environment.
Specialist AI Experience:
• Candidates should demonstrate practical experience in one or more of the following:
• Agent orchestration frameworks such as LangGraph, AutoGen, AWS Bedrock Agents or Microsoft Foundry Agent Service.
• Model Context Protocol, tool-calling APIs and agent state or memory management.
• Retrieval-augmented generation and vector database infrastructure.
• Cloud-agnostic model serving using tools such as ONNX, BentoML, Triton Inference Server or vLLM.
• MLOps platforms such as MLflow, Kubeflow or Airflow.
• GPU cluster management and inference or training workload optimisation.
• Prompt-injection prevention, output filtering, data-exfiltration controls and AI threat modelling.
•
AI Finops Experience:
• Candidates should have experience with some combination of:
• AI or cloud cost attribution and tagging.
• Chargeback and showback models.
• Token, GPU, DBU or provisioned-throughput cost management.
• Rightsizing, workload scheduling and reserved or spot-instance optimisation.
• Cost dashboards, anomaly detection and cost-per-use-case reporting.
• Communicating technical cost trade-offs to senior technology, business or finance stakeholders.
Qualifications:
• Postgraduate qualification in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering or a related quantitative field.
• A Master’s degree is preferred and may be required for certain senior appointments.
• Relevant certifications are strongly preferred, including:
o AWS Solutions Architect Professional or AWS Machine Learning
o Microsoft Azure AI Engineer
o FinOps Certified Practitioner
o Certified Cloud Security Professional or equivalent
o HashiCorp Terraform Associate
Kubernetes certification