תיאור המשרה
Meta is seeking a Staff Systems Software Engineer to design and build the foundational infrastructure that powers products used by billions of people worldwide. In this role, you will architect and implement large-scale distributed systems, low-level platform components, and high-performance services that underpin Meta's core product stack. You will drive technical strategy across system reliability, performance, and scalability, partnering closely with product, infrastructure, and data engineering teams to deliver systems that operate at global scale with high availability and efficiency.
Responsibilities
Architect and implement large-scale distributed systems and platform services that support high-throughput, low-latency workloads across Meta's product infrastructure
Lead the technical design of systems components including storage layers, compute pipelines, networking abstractions, and service orchestration frameworks
Identify and resolve systemic performance bottlenecks through instrumentation, profiling, and targeted optimization across the full systems stack
Define and enforce service level objectives for owned systems, building dashboards, alerting pipelines, and runbooks to reduce mean time to mitigation during incidents
Drive reliability improvements by reducing failure surface, designing resilient rollout strategies, and leading regular resiliency and overload testing exercises
Collaborate with cross-functional partners across product engineering, infrastructure, and data science to align system architecture with evolving product and business requirements
Establish and evolve coding standards, architectural patterns, and engineering best practices for systems development across the broader organization
Leverage AI-assisted development workflows to accelerate design iteration, code generation, and systems analysis, applying sound judgment on when to rely on AI versus deep systems expertise
Mentor other engineers on systems design principles, debugging methodologies, and production operations, and contribute to onboarding programs for new team members
Lead incident retrospectives, identify root causes of complex production failures, and drive implementation of systemic improvements to prevent recurrence
Minimum Qualifications
8+ years of experience designing and implementing large-scale distributed systems, platform infrastructure, or systems software in production environments
Experience leading major technical initiatives end-to-end, including architecture design, cross-team coordination, staged rollout, and post-launch reliability ownership
Experience debugging complex, non-reproducible systems issues including concurrency bugs, memory management failures, and distributed consistency problems
Experience defining service level objectives, building observability infrastructure, and driving reliability improvements across production systems
Experience communicating technical architecture decisions and trade-offs in writing to both engineering and non-engineering stakeholders
Preferred Qualifications
Experience with systems programming languages such as C, C++, or Rust in the context of high-performance or low-latency infrastructure
Experience building or improving developer tooling, automation frameworks, or internal platforms that measurably improve engineering efficiency across teams
Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Experience designing or contributing to storage systems, compute scheduling frameworks, or inter-service communication protocols at scale
Demonstrated use of AI tools to accelerate systems design, automate operational workflows, or improve code quality and test coverage