Developer Experience Software / Systems Reliability Engineer
- Austin, TX / Santa Clara, CA · Hybrid
- Full Time
- USD 140,000–300,000 / year
Nuvacore is building a ground-up, high-performance, low-power CPU for next-generation agentic AI compute workloads. We are seeking a Software or Systems Reliability Engineer with a strong background in developer tools/experience to build and own the front-end tool flows that our design and verification teams rely on to develop and verify the CPU. You will work closely with hardware design and verification engineers to keep build and test workflows (simulation, emulation, synthesis, lint, CDC, and formal verification) fast, reliable, scalable, and delightful.
THE ROLE
-
Build System Architecture: Design and own the build system that turns a large, hierarchical RTL and verification codebase into reproducible, incremental, cacheable builds. We use Buck2; you'll own the rule set, the dependency model, and the remote-execution and caching strategy as the design grows.
-
Continuous Integration: Design the pipelines that gate design changes, with a particular focus on build and test selection, so each change runs the smallest set of long-running tests that can catch its regressions.
-
Job Scheduling & Orchestration: Model and schedule heterogeneous, long-running, resource-constrained jobs (hours-long, memory-heavy, license-gated) across the compute farm, and build the orchestration layer that sequences multi-stage workflows and surfaces their results.
-
Distributed Storage & Data Management: Decide how build artifacts, test outputs, and design data are stored, cached, versioned, and expired across distributed storage, including content-addressed caches, retention policies, and data locality on the farm.
-
Reliability & Observability: Instrument the flows with metrics, logs, and traces; define and track SLOs for turnaround time, flakiness, and failure attribution; drive root-cause fixes rather than reruns.
-
Flow Integration: Wrap vendor tools as hermetic build actions with explicit inputs and outputs. Partner with design and verification engineers to understand their workflows.
REQUIREMENTS — MUST HAVE
-
BS in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
-
Experience designing and implementing complex build systems with hierarchical dependencies: dependency graphs, incrementality, hermeticity, and caching.
-
Hands-on experience with Bazel and/or Buck, specifically writing rules and macros.
-
Experience designing continuous integration systems, especially dependency-aware build and test selection.
-
Experience with job scheduling for long-running, resource-constrained workloads (Slurm, LSF, Kubernetes, or similar).
-
Experience with distributed storage systems and the tradeoffs they impose on consistency, locality, caching, and retention.
-
Strong software engineering in Python, Rust, Go, and/or shell: tested, reviewed, maintainable code, with infrastructure treated as software.
-
Solid understanding of Linux operation system administration.
REQUIREMENTS — NICE TO HAVE
-
Experience in a semiconductor or hardware design environment, building or supporting front-end EDA flows (RTL simulation, emulation, lint/CDC, synthesis, or formal).
-
Hands-on experience with EDA verification tools such as Synopsys VCS, Cadence Xcelium, and/or Jasper formal verification tools.
-
Hands-on experience with EDA emulation platforms (Synopsys Zebu, Cadence Palladium)
-
Working knowledge of RTL design and verification flows, enough to debug flow issues alongside design and verification engineers.
-
Experience with remote execution and remote caching (Buck2 RE, Bazel RBE, BuildBarn, BuildBuddy, NativeLink, or similar).
-
Experience with workflow orchestration engines (Nextflow, Snakemake, Airflow, Temporal, or similar).
-
Exposure to cloud compute for large engineering workloads (AWS, Azure, or GCP).
-
Familiarity with EDA tool licensing models and how they constrain scheduling (FlexLM/FlexNet).