Job Title: Senior Software Engineer (Agent SDK & Developer Enablement)
Job Location: Oaks,PA
Job Type: Contract
Job Description:
Design and build the agent SDK across multiple languages, published to internal package registries, wiring model calls through the platform gateway together with tracing, identity, guardrails, and retrieval and memory interfaces.
Own the public interface of the SDK: API design, versioning and backward compatibility, deprecation strategy, and the release process teams depend on.
Build the SDK interfaces agents use for memory and retrieval: session and short-term conversation state, long-term memory, and retrieval against the enterprise knowledge sources, exposed so that an agent developer gets correct behaviour without hand-rolling their own integration.
Build the project scaffold that generates a compliant, runnable agent project, with the SDK pre-wired, configuration, container and deployment manifests, evaluation configuration and the pipeline already in place.
Build the evaluation suite and quality baseline: evaluation criteria such as groundedness, safety, refusal behaviour and task correctness, seed datasets, and a standard test format that teams extend with their own cases.
Build the test and evaluation tooling agent developers actually use: harnesses for exercising agent behaviour, mocking model and tool calls for deterministic tests, replaying recorded runs, and asserting on non-deterministic output.
Deliver the CI/CD pipeline template shipped inside the scaffold, including the evaluation gate that blocks promotion when quality or safety thresholds are not met.
Build reference and template agents that prove the full path end to end, and a working template in each supported language for teams to clone.
Write the golden-path documentation, structured to be consumed both by engineers and by AI coding tools, and maintained as the platform changes.
Support onboarding of existing agents, providing documented migration steps and worked references for teams wrapping already-built agents in the SDK.
Work directly with the first consuming teams, turning their friction into changes in the SDK, the scaffold and the documentation.
3+ years hands-on agent or LLM application development, having personally built and shipped applications using agent frameworks or SDKs such as LangChain, LangGraph, the OpenAI Agents SDK, Microsoft Agent Framework, or equivalent. This is non-negotiable for this position.
6+ years professional software engineering, with production depth in at least two of C#, Python, Java and TypeScript or JavaScript, and the ability to work across all four. There is no one-language-per-engineer model on this team.
2+ years building automated testing or evaluation for LLM or agent applications, using frameworks such as the Azure AI Evaluation SDK, Promptfoo, DeepEval, Ragas, LangSmith evaluation or equivalent. This includes handling non-deterministic output, building golden datasets, and running evaluation as a gate inside CI rather than as a manual exercise.
2+ years working with agent memory and retrieval in application code: conversation and session state, short-term and long-term memory, and retrieval-augmented generation against a vector or search backend. You do not need to build the stores, but you must know what an agent developer needs from them and how these interfaces behave inside an agent loop.
3+ years designing and maintaining libraries or SDKs consumed by other engineering teams: public API design, versioning, backward compatibility, and packaging to registries such as NuGet, PyPI, Maven or npm.
3+ years building CI/CD pipelines using GitHub Actions, Azure DevOps or equivalent, including quality gates that block promotion.
2+ years working with LLM application patterns in production: tool and function calling, structured output, streaming, context management, retries and fallback across providers, and prompt handling.
Fluency with AI-assisted development tooling such as GitHub Copilot, Claude Code or Cursor, used daily as part of how you work. Team sizing on this engagement assumes that leverage.
2+ years working with containerised deployment and a major cloud platform, sufficient to build scaffolds that deploy cleanly into Kubernetes and managed runtimes.
Demonstrable developer experience judgement: building tooling, templates or libraries that other engineers adopted willingly, and iterating based on their feedback.
Working knowledge of OpenTelemetry and how tracing is instrumented inside a client library.
Experience with red teaming or adversarial testing of LLM applications, including prompt injection and jailbreak testing.
Experience building internal developer platforms or paved-path tooling adopted across multiple engineering teams.
Model Context Protocol (MCP) server or tool integration.
Experience with project scaffolding or code generation tooling, such as Yeoman, Cookiecutter, Backstage software templates or equivalent.
Technical writing for engineering audiences, including documentation structured for consumption by AI coding tools.
Experience in financial services or another regulated industry, where governance is expected to be built into tooling rather than applied afterwards.
Exposure to Azure AI Foundry or comparable managed agent runtimes.
By continuing you agree to our Terms & Privacy Policy.