
Staff Platform Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Create golden roads and paved paths utilizing reference implementations, shared libraries, service templates, scaffolding, and pipeline templates.
• Assess the adoption of paved paths and the time taken to achieve initial success, while enhancing areas that teams typically navigate around.
• Execute the first genuine production workload with an actual team.
• Develop distributed-systems patterns that include backend services, event-driven workers, job pipelines, APIs, and relational data models.
• Collaborate with Platform Engineering to construct, plan, and transfer shared infrastructure.
• Oversee the migration and adoption of existing services to new paths and phase out replaced systems.
• Engage in cross-team code reviews and respond to production incidents.
• Define and uphold target-state architecture alongside owners and decision timelines.
• Manage AWS foundations for platform and AI workloads, covering accounts, identity, networks, isolation, runtimes, managed data services, tagging, and cost governance.
• Promote standards adoption through tooling and pipelines.
• Lead decisions regarding buy/build/reuse for models, frameworks, and vendors.
• Integrate overlapping tools and standards while managing decommissioning strategies.
• Design and develop production-grade AI agent systems, multi-agent workflows, orchestration layers, and supporting services.
• Create retrieval, context, and memory systems that incorporate sensitive-data scoping, redaction, and auditing.
• Develop AI decision-support systems that meet the accuracy, traceability, and explainability requirements of healthcare and regulated environments.
• Construct React and TypeScript user experiences for AI functionalities.
• Establish AI quality-assurance practices, including automated evaluations, regression testing, groundedness checks, and compliance guardrails.
• Create a standardized approach to building, running, evaluating, and operating AI systems.
• Connect teams, reveal duplicated efforts, assign unowned decisions, and synchronize cross-team initiatives.
• Foster communication through writing, demonstrations, and working sessions.
• Contribute to roadmap and capacity planning efforts.
• Facilitate collaboration among Engineering, Product, Data, Security, and Infrastructure teams.
• Represent technical direction to leadership and convey business context to engineering teams.
• Collaborate with Product to translate vague business challenges into architectural solutions.
• Enhance engineering standards through design reviews, mentorship, and decision records.
• Work with Security, Compliance, and Data teams on AI governance, access control, and data management.
• A minimum of eight years of software engineering experience, encompassing architecture scope, organization-wide responsibilities, and ownership of complex systems from design through production.
• Current hands-on coding and code review experience.
• Experience in building internal platform capabilities that are voluntarily adopted by other teams.
• Production experience with AI applications, including evaluation, monitoring, failure management, guardrails, latency, and cost control.
• Extensive AWS expertise across multi-account environments, identity and permission boundaries, networking, compute and serverless runtimes, managed data services, managed model services, and cost/tagging governance.
• Familiarity with Infrastructure as code using Terraform, CDK, or both.
• Proven track record in establishing and promoting architectural foundations and standards.
• Practical experience with agentic software development utilizing AI coding agents.
• Profound knowledge in Python or JVM backend and distributed systems, or React and TypeScript product engineering.
• Capability to work across the entire system.
• Experience with retrieval, embeddings, context, memory, evaluation, monitoring, or AI agent architectures and workflows.
• Strong judgment in testing, monitoring, failure management, privacy, security, access control, and data governance in regulated or audited environments.
• Experience collaborating with platform or infrastructure teams to implement organization-wide changes.
• Ability to serve as a central technical point across multiple teams without acting as an approval gate.
• Demonstrated capacity to align conflicting teams and contribute to roadmap and capacity planning.
• Excellent written communication and effective collaboration with both technical and non-technical teams.
• United States work authorization is required; a visa sponsorship inquiry is included in the application process.
• Preferred: experience in healthcare, health insurance, or benefits administration, including ICHRA, claims, eligibility/enrollment, EDI, or FHIR.
• Preferred: experience with a venture-backed high-growth company.
• Preferred: work related to internal developer platforms or developer experience.
• Preferred: experience with GitLab CI at scale.
• Preferred: experience in production observability and LLM observability, including Datadog.
• Preferred: familiarity with event-driven and streaming architectures such as Kafka or Flink.
• Preferred: experience with agent tooling and standards such as Claude Code, MCP, and agent SDKs.
• Preferred: experience with AI governance or technology radar in collaboration with security or compliance partners.
• Coverage for alternative medicine.
• Flexible paid time off (PTO).
• Up to 16 weeks of paid parental leave.
• Paid holidays.
• 401(k) retirement plan.
• Transportation benefits.
• Education reimbursement options.
• Two days of paid paw-ternity leave.
• Standard health and wellness benefits.
• Equity options available.
Planet Technologies
Northrop Grumman
Credit Acceptance
Planet Technologies
Get handpicked remote jobs straight to your inbox weekly.