
Senior System Software Engineer β Scientific Computing PaaS
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in California.
β’ Take charge of designing services and managing the underlying cloud infrastructure for scientific workflows that are informed by physics and driven by data.
β’ Develop innovative algorithms and collaborate with operations to enhance the overall performance of the system across the entire stack.
β’ Engage with application code, which includes deep learning frameworks, numerical solvers, microservices, APIs, and heterogeneous CPU/GPU computing.
β’ Create, deploy, and manage scalable I/O infrastructure for tasks such as checkpointing, data loading, and pre- and post-processing of data.
β’ Improve the compute, storage, and network architecture for applications based on physics and simulations.
β’ Construct scientific computing platform workflows on the cloud for numerical simulation solvers, AI training, inference, and visualization.
β’ Provide support for applications related to weather forecasting, climate modeling, industrial design, and digital twin simulation across various domains.
β’ A BS/MS degree in Computer Science or a related field, or equivalent professional experience.
β’ Over 10 years of experience in creating and managing distributed compute and data-intensive platforms as a service on the cloud.
β’ Demonstrated proficiency in a compiled programming language such as Go, Rust, C++, or similar.
β’ Solid foundational understanding of Cloud Computing, including data center architecture, cloud security, virtualization, resource pooling, and elasticity.
β’ Proven expertise in Distributed Systems and Parallel Processing, encompassing distributed computation models, synchronization, deadlock detection, fault tolerance, consensus protocols, parallel algorithms, and shared and distributed memory architecture.
β’ Practical debugging experience with processes, threads, deadlocks, synchronization, scheduling, IPC, memory management, file systems, and I/O structures.
β’ Strong skills in algorithmic thinking and system design, including recursion, graphs, trees, stacks, queues, and the design and operation of large-scale loosely coupled distributed systems.
β’ Self-motivated with robust interpersonal skills, capable of working independently and effectively with multiple teams with minimal supervision.
β’ [Preferred/standout] Experience in building, deploying, and managing AI platforms on HPC clusters and cloud-native systems.
β’ [Preferred/standout] Experience with distributed storage, scheduling, orchestration, as well as compute, storage, and network systems.
β’ [Preferred/standout] Experience in configuring and troubleshooting hardware, operating systems, kernels, and compilers to achieve maximum performance.
β’ [Preferred/standout] Hands-on debugging experience to optimize compute, networking, and I/O frameworks.
β’ [Preferred/standout] Experience in debugging and customizing third-party source code.
β’ Equity
β’ Comprehensive benefits package
NVIDIA
FXC Intelligence
Miratech
FXC Intelligence
Get handpicked remote jobs straight to your inbox weekly.