Job Description: AWS Cloud Platform & DevOps Engineer (GenAI Platform) - Contractor
Join us as an AWS Cloud Platform & DevOps Engineer (GenAI Platform) - Contractor within the IB Analytics team. The role owner will be responsible for hands-on design, provisioning, automation, deployment, monitoring, and production-readiness of the AWS platform supporting AI agents, MCP services, skills, orchestration components, and data integrations. The successful candidate will bring deep practical expertise across AWS, ECS, EKS, Kubernetes, CloudFormation, IAM, CI/CD pipelines, SonarQube, ESAAS, observability, and secure deployment automation.
Key Responsibilities
Design, provision, enhance, and maintain AWS environments supporting core platform services, agents, MCP servers, skills, and orchestration components.
Build and configure ECS and EKS platforms, including Kubernetes deployment standards, ingress, load balancing, scaling, namespace design, and environment separation.
Automate infrastructure provisioning using CloudFormation, AWS CDK, Terraform, or equivalent infrastructure-as-code tooling.
Configure VPCs, subnets, route tables, ACLs, security groups, API Gateway, load balancers, S3, RDS/PostgreSQL, ElastiCache, Lambda, Systems Manager, Parameter Store, and Secrets Manager.
Design and implement IAM roles, permission boundaries, service roles, system accounts, and secure cross-service access patterns.
Build and manage CI/CD pipelines for BankerOne core platform, agents, MCPs, skills, and supporting services across non-production and production environments.
Integrate automated unit-test checks, test coverage, code quality gates, SonarQube, static scanning, security checks, and deployment governance into delivery pipelines.
Implement controlled deployment approaches where appropriate for platform and application services.
Implement CloudWatch dashboards, logging pipelines, alerting, tracing, performance monitoring, availability monitoring, and ESAAS integration.
Support security validation, resilience checks, production readiness, runbook creation, knowledge transfer, and RTB handover activities.
Essential / Basic Qualifications
To be successful in this role, you should have:
Strong hands-on experience in AWS cloud platform engineering, DevOps, or site reliability engineering.
Deep expertise with ECS, EKS, Kubernetes, Docker, container networking, orchestration, scaling, and deployment patterns.
Proven experience with CloudFormation and infrastructure-as-code automation; experience with AWS CDK or Terraform is highly relevant.
Strong working knowledge of AWS networking including VPCs, subnets, routing, ACLs, security groups, API Gateway, and load balancers.
Strong knowledge of IAM, service roles, permission boundaries, Secrets Manager, Systems Manager, Parameter Store, and secure credential management.
Experience with PostgreSQL/RDS, S3, ElastiCache/Redis, Lambda, CloudWatch, API Gateway, and related AWS managed services.
Strong CI/CD experience using Jenkins, GitLab CI/CD, GitHub Actions, Harness, Azure DevOps, or similar tooling.
Experience implementing SonarQube, static code analysis, security scanning, automated test checks, and quality gates.
Experience implementing logging, monitoring, alerting, dashboards, and operational telemetry for production services.
Ability to independently troubleshoot complex deployment, networking, IAM, Kubernetes, pipeline, runtime, and observability issues.
Experience working in globally distributed agile teams with strong ownership and delivery accountability.
Bachelor degree in a technical discipline such as Computer Science, Engineering, or equivalent experience.
Desirable / Good to Have Skills
Experience deploying GenAI, LLM, agentic AI, MCP, or data-intensive workloads on AWS.
Working knowledge of AWS Bedrock, model-hosted application patterns, inference workloads, or AI platform services.
Experience with OpenTelemetry, distributed tracing, structured logging, and enterprise observability platforms.
Experience with blue-green deployment, canary deployment, automated rollback, and controlled production release patterns.
Familiarity with enterprise security, architecture, cloud adoption, governance, and production-readiness review processes.
Experience supporting RTB handover, ResCat validation, runbook creation, service monitoring, and operational support documentation.
Financial services, Investment Banking, front-office technology, or other regulated enterprise technology experience.
Exposure to SonarQube, ESAAS, CyberArk, service accounts, secrets rotation, cost management, and cloud governance tools.
Other Skills & Attributes
Strong problem-solving mindset with the ability to debug complex cloud, network, deployment, and pipeline issues.
Excellent communication skills and the ability to collaborate with application developers, architects, security teams, governance teams, and production support.
Independent, reliable, and highly delivery-focused with strong ownership of platform readiness.
Comfortable operating in a fast-paced, enterprise-controlled environment with multiple governance dependencies.
Willingness to define reusable platform patterns, improve engineering standards, and contribute to team knowledge sharing.