Gain full access to exclusive job listings from leading companies worldwide.
Verified, High-Quality Jobs Only
No ads, scams, or junk-just genuine opportunities.
Focus on Real Opportunities
Explore thousands of open positions tailored to your lifestyle, including flexible remote jobs.
Exclusive Resume Review
Receive expert feedback with personalized suggestions to enhance your resume.
This role is responsible for driving the reliability, resiliency, performance, and modernization of critical American Express platforms across Distributed environments. You will leverage deep technical expertise in software engineering, runtime engineering, production support, and platform operations to quickly assess and remediate complex availability, performance, and operational issues. As part of our technology team, you will partner with engineering, product, infrastructure, and operations teams to design, build, automate, and support highly available enterprise platforms. You will help accelerate modernization initiatives, improve operational excellence, and deliver secure, scalable, and resilient solutions that power critical customer and business capabilities.
Software Engineering & Platform Development
Serve as a hands-on engineer with experience in supporting complex enterprise applications, platforms, and operational tooling across Distributed environments.
Design, develop, prototype, code, test, and implement scalable software solutions using technologies such as Java, Python, SQL, and related frameworks.
Act as a technical contributor in change management reviews, root cause analysis, and troubleshooting of complex technical issues.
Design and implement automation solutions, and engineering practices that improve platform resiliency, operational efficiency, and security.
Use best practices in incident management, problem management and change management as this role focuses on application production support.
Runtime Engineering, Reliability & Operations
Contribute to the technical roadmap for runtime systems, ensuring platform reliability, scalability, availability, recoverability, and performance.
Establish, monitor, and continuously improve key performance indicators (KPIs), service level objectives (SLOs), and operational metrics (MTTR, MTBF) for platform health and resiliency.
Perform diagnosis and resolution of production incidents, batch failures, application outages, performance bottlenecks, and infrastructure issues across Mainframe and Distributed platforms.
Apply Site Reliability Engineering (SRE) principles and operational excellence practices to improve system stability and reduce operational risk.
Support disaster recovery, high availability, workload management, capacity planning, and business continuity initiatives.
Distributed Platform Engineering
Must have experience in Support and optimize enterprise platforms across distributed technologies including Java, Python, C, SQL, NodeJS, Bash, JS/HTML/CSS, Spanner, BigTable, BigQuery, Spring, GraphQL, OOP, MVC, Algos, Git, CoPilot, Jenkins, XLR, Elastic, Jira, AI, ML, Docker, Kubernetes, Kafka, Rest API, GRPC, Grafana.
Desirable Support and optimize Distributed Platform technologies including cloud infrastructure, Linux/Unix, containers, APIs, Java-based services, distributed databases, and modern application platforms.
Implement and support Cloud/Distributed architectures, modernization initiatives, API enablement, and enterprise integration capabilities.
Knowledge of cloud platforms AWS, GCP or general Cloud fundamentals.
Data, Integration & Automation
Develop and support enterprise integration solutions utilizing APIs, MQ, Connect:Direct, event-driven architectures, and batch and real-time data integration patterns.
Utilize relational and NoSQL databases including DB2, PostgreSQL, Redis, and Couchbase to support critical business applications.
Automate operational processes, deployments, monitoring, reporting, and remediation activities using Python, Bash, Ansible, Jenkins, and related technologies.
Collaborate with engineering teams to adopt scalable automation and self-service capabilities for deployment, monitoring, and operational support.
Observability & Continuous Improvement
Implement and utilize monitoring, observability, logging, and analytics solutions using tools such as Splunk, OMEGAMON, RMF/SMF, Sysview, MainView, Dynatrace, AppDynamics, ELK, or equivalent technologies.
Analyze operational trends, identify opportunities for optimization, and formulate strategic recommendations to improve platform health and engineering effectiveness.
Contribute to continuous improvement initiatives focused on reliability, performance, security, operational maturity, and customer experience.
DevOps, Security & Governance
Understand CI/CD pipelines and DevOps practices using tools such as Git, Jenkins, Maven, DBB, Endevor, Changeman, ISPW, UrbanCode Deploy, or equivalent platforms.
Apply enterprise security controls, compliance requirements, audit standards, and access management practices, including RACF, ACF2, Top Secret, and cloud security principles.
Ensure solutions meet non-functional requirements (NFRs) including availability, scalability, performance, security, recoverability, and maintainability.
Bachelor's degree in Computer Science, Computer Engineering, and/or comparable experience
Work experience in software engineering, app support or infrastructure operations or runtime engineering.
A working understanding of cloud infrastructure, distributed systems, and containerization technologies, with experience in supporting critical business applications being a plus.
Familiarity with monitoring and logging tools, and incident management best practices, to ensure reliability and performance of applications in a production environment.
Solid programming and scripting skills, with hands on experience to automate operational tasks using tools such as Python
Knowledge of scripting languages (e.g., PowerShell, Python) for automation tasks
Experience in technology operations work
Hands on experience with relational and NoSQL databases such as DB2, Redis, Postgres, Couchbase etc.
Experience in cloud platforms such as AWS, Azure, or Google Cloud, Public Cloud certification is a plus
At American Express, our culture is built on a 175-year history of innovation, shared values and Leadership Behaviors, and an unwavering commitment to back our customers, communities, and colleagues. From delivering differentiated products to providing world-class customer service, we operate with a strong risk mindset, ensuring we continue to uphold our brand promise of trust, security, and service.
As part of Team Amex, you’ll experience our powerful backing with comprehensive support for your holistic well-being and many opportunities to learn new skills, develop as a leader, and grow your career. Here, your voice and ideas matter, your work makes an impact, and together, you will help us define the future of American Express.