We are on the lookout for a Principal Software Engineer to join our Elasticsearch - Distributed Systems team and focus on how Elasticsearch provides scale, performance, and resilience, especially in the context of our Serverless offering
You will be tackling deep distributed systems challenges in how nodes in an Elasticsearch cluster communicate, and how data are indexed, allocated, and managed across nodes
Leveraging your deep distributed systems expertise to design, build, and operate core pieces of our next-generation Serverless Elasticsearch offering, directly impacting the experience of customers worldwide
Improving Elasticsearch’s components that support concurrent and consistent indexing across multiple machines
Maintaining our cluster coordination system to keep performance high even though nodes come and go from the cluster and data moves around, while maintaining the safety and liveness properties of the system as a whole
Pushing the limits on the number of shards, nodes, and petabytes that Elasticsearch can handle today
Looking into all kinds of issues, including performance or concurrency issues, and proposing solutions
Supporting our support engineers with the harder problems
Owning operational investigations end-to-end, until performance or stability issues are successfully resolved
Benefits
Toast to your health: Fully paid health coverage for you and your family, in many locations.
Craft your calendar: Flexible location and schedule for most roles.
Create space for you: Distributed by design workforce, plus generous number of vacation days each year.
Embrace parenthood: Minimum of 16 weeks of parental leave, plus generous family formation benefits.
Give back your time: 40 hours each year to use toward volunteering with organizations and causes you’re passionate about.
Amplify your impact: Double your charitable giving — we match donations up to $1500 USD (or local currency equivalent).
You have experience building large-scale distributed systems, rather than making use of off-the-shelf services
You have a deep background in distributed systems and consensus algorithms
You have experience managing projects involving multiple engineers
You have a deep technical proficiency in algorithms
You have a proven track record of using AI to accelerate development, debug complex systems, and optimize code, while still owning the final outcomes
You possess the ability to collaborate across functions and teams and seamlessly transition between different projects, codebases, or teams based on business priorities
You have strong skills in core Java and are conversant in the standard library of data structures and concurrency constructs, as well as newer language features
You are able to own projects from beginning to end. This covers both technical design and working with others to develop needed components
You have experience building and debugging complex features with a broad impact, running on multiple machines
You can work autonomously, drive decisions and results in a distributed team by leveraging asynchronous, direct, and transparent communication
Asynchronous event-driven network frameworks such as Netty