Skip to main content

Senior Software Engineer, Infrastructure & Systems

Astronomer

  • New York City
  • Remote
  • fulltime
  • Posted today

Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data. To learn more, visit www.astronomer.io .

About This Role

Senior and Staff Software Engineers on our Infrastructure team own the architecture of the systems that stand up, configure, and manage the infrastructure the underlying Airflows run on. This includes the security of those systems and the reliability and observability capabilities that keep Astro running at scale, across both Astronomer's multi-tenant cloud platform and Astro Private Cloud, our self-hosted distribution for regulated organizations running Airflow in their own private clouds or fully isolated data centers.

You'll design and build the control-plane services, the APIs that drive infrastructure lifecycle management, and the control/data-plane interaction patterns that make it possible to bring up and manage that infrastructure — including how those interactions hold up across network topologies from standard private cloud to zero-egress, air-gapped environments.

You'll bring deep, hands-on expertise in distributed systems, API design, and infrastructure lifecycle management, own the technical direction, design, and implementation of key areas of the platform, and mentor engineers earlier in their careers. This posting covers both Senior Software Engineer and Staff Software Engineer, with the level you're hired at depending on experience and demonstrated scope.

What You'll Do

  • Design and own the architecture of the systems that stand up, scale, and manage the infrastructure Airflow runs on, leading these efforts end-to-end (design, implementation, tests, documentation, and rollout).
  • Design and evolve the APIs (REST/gRPC) that drive infrastructure lifecycle management, powering customer-facing UI and programmatic integrations, including data modeling, versioning, and backward compatibility.
  • Own the networking and security posture of these systems, including how they authenticate, authorize, and communicate securely, and drive CVE remediation, image hardening, and Pod Security Standards compliance.
  • Architect observability and traceability across the stack: metrics, logs, and traces that let a request be followed end-to-end across service and network boundaries.
  • Participate in on-call rotation, contributing to incident diagnosis, resolution, and post-mortems for these systems and ensuring remediation items are tracked into sprints.
  • Mentor engineers and raise the bar on design and code review, testing, and operational rigor.

What You'll Bring

  • 5+ years of experience in infrastructure, platform, or systems engineering, with a track record of owning production systems at scale (8+ years typical for Staff).
  • Deep, hands-on expertise with Kubernetes, including building custom Operators/CRDs and managing Helm-based deployments for infrastructure lifecycle workflows.
  • Experience designing APIs (REST or gRPC) that serve as the core engine behind a customer-facing UI or programmatic integrations, including versioning and backward compatibility.
  • Strong systems fundamentals: networking, distributed systems, reliability engineering, and security best practices.
  • Proficiency in Go and/or TypeScript, or a demonstrated ability to pick up new languages quickly.
  • Practical experience designing and/or building observability and distributed tracing for multi-layer systems
  • Demonstrated ability to lead technical design across teams and communicate trade-offs to both engineers and stakeholders.
  • Experience mentoring engineers and improving team-wide practices.

Nice to Have

  • Experience designing or operating multi-tenant control planes.
  • Experience shipping software into air-gapped or highly regulated environments (finance, healthcare, government).
  • Familiarity with Apache Airflow internals or other data-orchestration platforms.
  • Incident command / on-call leadership experience.

The estimated total compensation for this role ranges from $200,000 - $300,000 based on leveling and geography, along with an equity component and a comprehensive benefits package. This range is merely an estimate; actual compensation may deviate from this range based on skills, experience, and qualifications.

At Astronomer, we value diversity. We are an equal opportunity employer: we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.