Create Alert
Email me similar jobs

Cloud Platform Engineer (Multi-Cloud & Hybrid Infrastructure)

Contract Latency DNS Tokens Azure Runbook 70 USD

Job Title: Senior Cloud Platform Engineer (Multi-Cloud & Hybrid Infrastructure)
Job Location: Oaks,PA
Job Type: Contract

Job Description:

Build and operate the Kubernetes runtime for agents, including namespace design, workload identity, ingress, autoscaling and cluster security, delivered as infrastructure-as-code.

Design and implement connectivity across the estate: cross-cloud networking, VPC and virtual network peering, private endpoints, private DNS resolution across boundaries, and NAT and egress control.

Establish runtime patterns on hyperscalers beyond the primary cloud, so agents deployed there use the same deployment path, identity model and telemetry contract as everywhere else.

Deliver private connectivity to on-premises infrastructure, including the paths agents use to reach an on-premises GPU cluster, with latency validation and the firewall and egress allowances telemetry depends on.

Build the infrastructure AI workloads depend on: container registries and image supply chain, secrets and key management, service-to-service identity, and the capacity and quota model for inference endpoints.

Establish the network paths for third-party and self-hosted platform services, such as observability backends, model providers outside the primary cloud, and enterprise systems the platform integrates with. This covers private connectivity or controlled egress, DNS resolution, certificate and TLS handling, and the firewall allowances each integration needs.

Own the infrastructure-as-code estate: module design, state management, and promotion from lower environments through to production across every runtime target.

Work with client security, network and infrastructure teams to take changes through review, including data residency, egress and network segmentation considerations.

Produce the topology documentation, architecture decision records and runbooks that the wider team and the client operate from.

MUST HAVE

7+ years hands-on cloud infrastructure engineering in production environments, covering compute, networking, identity, storage and monitoring.

Production depth on at least two major hyperscalers (Azure, AWS or Google Cloud), including the networking and identity model of each. Breadth here is a defining requirement, not a bonus: you will be expected to design across cloud boundaries rather than within one.

5+ years Kubernetes in production, including self-operated or unmanaged clusters. Cluster networking, ingress, workload identity, autoscaling and cluster security, not managed control planes alone.

4+ years enterprise cloud networking: virtual networks and VPCs, subnets, peering, security groups, private endpoints, private DNS resolution and cross-network routing.

3+ years hybrid cloud and on-premises connectivity, using ExpressRoute, Direct Connect, site-to-site VPN or equivalent, including DNS resolution across boundaries and firewall and egress rules. On-premises infrastructure is a first-class runtime target on this platform, not an edge case.

3+ years integrating third-party or self-hosted services into an enterprise network, including private connectivity or controlled egress, DNS resolution, TLS and certificate management, authentication between systems, and working the firewall and security review needed to get each path approved.

4+ years infrastructure-as-code to production standard using Terraform or equivalent, including module design, state management and multi-environment promotion.

3+ years building infrastructure for AI or data-intensive workloads, on any major cloud: inference or model endpoints, container registries and image supply chain, service-to-service identity, secrets management, and the capacity and quota model these workloads run under.

Working knowledge of the AI service landscape across hyperscalers: Azure AI Foundry and Azure OpenAI, AWS Bedrock, Google Vertex AI or equivalent, and an understanding of what differs between them in networking, identity and cost.

2+ years building CI/CD pipelines for infrastructure delivery using GitHub Actions, Azure DevOps or equivalent.

Working knowledge of LLM platform operations: model endpoints, token throughput, quotas and rate limits, and how inference cost accrues and is attributed.

PREFERRED

Cloud certification on more than one platform, such as AWS Solutions Architect Professional, Google Professional Cloud Architect or Azure Solutions Architect Expert.

Certified Kubernetes Administrator (CKA) or comparable depth demonstrated in production.

Experience operating self-hosted or GPU-based inference infrastructure, including capacity and throughput planning for on-premises model serving.

Experience in financial services or another regulated industry, where security review governs the pace of environment change.

Service mesh, private link or zero-trust networking patterns applied across cloud boundaries.

Familiarity with OpenTelemetry or platform observability tooling, particularly telemetry egress from constrained environments.

Exposure to AI gateway or model-routing patterns, such as gateway policies applied to model traffic or token-based rate limiting.

Similar jobs

Cloud Platform Engineer (Multi-Cloud & Hybrid Infrastructure)

Apply Now
Back to search page