Acerca del rol
Role Overview
ATOM is an end-to-end insurance operating system for MGAs, brokers and insurers, powering elseco, a DFSA-regulated specialty MGA. We build full-stack applications on AWS and are now building a fully orchestrated AI platform: an AI-enabled ingestion and orchestration layer with agents working under human oversight, embedded as ATOM intellectual property.
This role designs, builds and operates the cloud and AI platform beneath those applications. It is a hands-on engineering role. The platform you build must be defined in code, secure by design, observable, resilient, cost-controlled and audit-ready. You will work closely with application engineering (Node.js), data, architecture, and the Infrastructure and Cyber Security managers.
Key Responsibilities
- Architect and build the AWS foundation for ATOM: accounts, networking, identity, encryption and shared services, all as reviewed, reproducible code.
- Build and operate the AI orchestration layer: workflow orchestration, Bedrock agents and retrieval, guardrails, evaluation pipelines and human-review queues, starting with the aviation submissions pilot.
- Enable application teams with secure, paved-road patterns for containers, functions, APIs and eventing, so engineers ship without bespoke infrastructure work.
- Own CI/CD design and operation: pipelines, environment promotion, policy checks, secrets handling and release controls.
- Instrument everything: metrics, logs, traces and AI-specific telemetry, with alerting that detects issues before users do.
- Control cost: tagging, budgets, model and capacity choices, and monthly unit-cost reporting with a clear optimisation backlog.
- Design and test DR/BC for the platform with the Infrastructure and Cyber Security managers, and drive action closure from each test.
- Maintain data-flow architecture from public entry points through application services, private networking and data stores, including which data reaches which model. Keep it current enough to use for incident triage, change impact analysis and DFSA or audit requests.
- Make the platform independent of individuals: runbooks, standards, documentation and knowledge transfer to the wider engineering team.
Requirements
- Proven design and operation of production AWS workloads across compute, orchestration, eventing, networking and storage.
- At least 8 years in software, platform or infrastructure engineering, including at least 4 years designing and operating production workloads with AWS as the primary platform, and demonstrable production experience running LLM or AI services.
- Expert infrastructure-as-code in Terraform or AWS CDK, with no manual changes in production.
- Production-standard programming in Python and/or TypeScript/Node.js.
- Hands-on production experience with Amazon Bedrock or an equivalent LLM platform (Azure OpenAI, Vertex AI), including retrieval-augmented generation, guardrails, evaluation and cost controls.
- CI/CD pipeline design with GitHub Actions or equivalent.
- Cloud security design: IAM, encryption, secrets management and private network patterns.
- Observability engineering across metrics, logs and distributed tracing.
- DR/BC
Experience
RTO/RPO, restore testing, DR runbooks and test execution.
- Solid Linux production operations.
- Ability to document data-flow architecture clearly for engineers, auditors and regulators.
- Comfort operating in a regulated, audit-driven environment, with strong documentation and runbook discipline.
Fuente: la propia página de carreras del empleador.