
Senior Infrastructure Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants anywhere in the world.
• Ensuring optimal performance and reliability. You will enhance our CI/CD pipelines to make them faster and more resilient, enabling smooth deployments and transforming each incident into a valuable lesson rather than a recurrence.
• Updating our infrastructure. While Buffer has a history, our infrastructure is always evolving. From deploying KEDA and Argo Rollouts to refining our monitoring’s signal-to-noise ratio, there is ample opportunity for improvement. Modernization isn't complete without AI; we are already utilizing it for investigations and boilerplate tasks, and aim to integrate it more deeply into our daily operations.
• Viewing engineers at Buffer as customers. Your role will involve developing the developer tools that streamline the inner loop, coupled with documentation that remains reliable at all hours or when accessed by an agent.
• Overseeing the daily reliability of our production platform. Maintain a stable environment for EKS, ArgoCD, and the AWS ecosystem, adjusting autoscaling to ensure the system responds effectively under load, and approaching incident response in a way that promotes learning from each incident instead of repeating mistakes. (On-call responsibilities are shared among all engineers at Buffer, with a week-long shift approximately once every quarter.)
• Establishing progressive delivery that earns the trust of the engineering team. Implement Argo Rollouts with clear rollback mechanisms, minimizing the time between identifying a problematic deployment and reverting it to mere seconds or minutes.
• Developing developer tools as products rather than scattered scripts. Enhance our in-house local development environment, BIBEs (Buffer Isolated Build Environments, full-stack staging deployments per PR), and our CLI tools to ensure a fast, frictionless, and parallel-friendly inner loop for AI agents. Track adoption, engage with users, and iterate based on feedback.
• Minimizing operational toil through AI. Automate low-risk workflows end-to-end so that the team can focus on challenging problems instead of repetitive tasks. AI will not directly manage infrastructure but will enhance the efficiency of those who do.
• Keeping our stack up to date. Lead lifecycle upgrades for application runtimes (Node.js, Python), Kubernetes, EKS, Helm versions, and the Terraform-managed areas, while also addressing security vulnerabilities on the infrastructure side.
• Enhancing the economic efficiency of our platform. Take charge of visibility efforts on Datadog, AWS rightsizing, and log filters to ensure that observability and cloud expenditure grow at a slower pace than the company itself.
• Collaborating with EPD on the platform they are building. Elevate the standard of documentation within the team, contribute to weekly security tasks (dependency and vulnerability management is a shared responsibility), and help the infrastructure team move towards shared ownership and reduced reliance on individual team members.
• You possess significant experience as an Infrastructure Engineer, SRE, "DevOps" engineer, or in a related role that qualifies you as a senior professional.
• You have practical experience managing production Kubernetes environments at scale on a managed service (GKE, EKS, AKS), including the creation and maintenance of Helm charts, and are proficient with autoscaling mechanisms such as KEDA and the cluster auto scaler.
• Your expertise in AWS includes IAM, EC2, S3, SQS, ECR, and ALBs. Familiarity with Cloudflare (WAF, Workers, etc.) and GCP (BigQuery) is a plus.
• You are adept at using Terraform, preferring to structure your code using modules while ensuring it remains adaptable, readable, and self-contained. Additional credit if you have contributed to an open-source module.
• You have experience operating production CI/CD with GitHub Actions (or alternatives) and GitOps using ArgoCD (or similar). You have created ArgoCD pipelines and Helm configurations, including canary or progressive delivery systems that you would trust to revert safely.
• You have developed internal developer tools (CLIs, development environments, per-PR environments) and approach them as products with users rather than mere scripts.
• You have a history of making pragmatic build-vs-buy decisions regarding infrastructure tools, able to justify your choices and revisit them as circumstances evolve.
• You have experience with observability tools like DataDog or Sentry, designing logs and metrics with cost considerations in mind, understanding that observability and cloud expenses can escalate if not monitored.
• You are comfortable working with Cloudflare, including Workers, Zero Trust, DNS, and other parts of their platform.
• You can read and modify TypeScript or Node services sufficiently to upgrade runtimes and assist teams (legacy PHP and Python experience is also beneficial).
• You are proficient with modern AI tools, utilizing them to debug, document, and minimize toil rather than solely generating code, and you apply these practices to how infrastructure operates.
• You are proactive and follow through on tasks. You identify necessary actions without prompting and ensure completion without needing reminders.
• You convert ambiguity into tangible concepts. You take vague requests, deliver initial drafts for team feedback, and iterate until the final product meets expectations.
• You excel in remote, asynchronous work environments, communicating clearly and providing ample context.
• You do not wait for perfect information to begin tasks, nor do you hold out for perfection to deliver results.
• You view infrastructure as a facilitator for engineering success rather than a barrier.
• You are invested in Buffer's customers. You feel their pain when performance lags. When facing flaky errors, you seek to identify root causes and aim to eliminate entire classes of issues, viewing the system holistically rather than just addressing individual bugs.
• You prioritize performance. If something is sluggish, it should not exist. You prefer to enhance performance rather than find workarounds.
• You are a versatile engineer with strong areas of expertise: T-shaped individuals with deep infrastructure knowledge and the adaptability to adjust as priorities shift.
• You create rather than just consume. Whether through open source contributions, technical blogging, conference presentations, side projects, or active engagement on platforms that Buffer supports, you showcase your creativity. We are a Team of Creators, and the closer infrastructure aligns with the creator's experience, the more effective the platform becomes.
• You consider infrastructure as a platform for its users. You have developed APIs, SDKs, CLIs, MCP servers, or developer-facing tools where user adoption, not just deployment, was the key success metric. You understand the difference between code that is merely shipped and code that is actively used.
• You think long-term. You prefer to invest in solid fundamentals rather than chase the latest trends.
• Bonus points if you are already a Buffer user or have familiarity with social media management tools.
• Offers Equity
Tailscale
Adzuna
Webflow
Coinbase
Get handpicked remote jobs straight to your inbox weekly.