ARCHIVED
This job listing has been archived and is no longer accepting applications.
MisuJob - AI Job Search Platform MisuJob

Member of Technical Staff - Site Infrastructure (US Government)

Xai

Los Angeles, CA; Memphis, TN (Los Angeles, CA, Memphis, TN, Palo Alto, CA) permanent

Posted: March 27, 2026

Interested in this position?

Create a free account to apply with AI-powered matching

Quick Summary

We are looking for a highly motivated and skilled engineer to join our team in Los Angeles or Memphis. The ideal candidate should have experience in Site Infrastructure and be familiar with AI technologies. Strong understanding of computer systems is a plus.

Job Description

About xAI

xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

You will be the person who turns a hardware listing and a software bundle into a running AI inference platform — from bare metal to serving production traffic. This is a hands-on role at the intersection of physical datacenter infrastructure and platform engineering. You will rack GPU servers, cable network fabrics, provision bare metal via PXE, deploy Kubernetes clusters, stand up monitoring and network telemetry stacks, and validate end-to-end inference pipelines — all in air-gapped, classified environments with no internet access.

You are the high side. Everything the platform engineering team builds on the unclassified side — deployment tooling, signed software bundles, switch configurations, OS images — you execute on classified infrastructure. You own the full stack from physical hardware through running GPU workloads, including the cross-domain solution (CDS) receive pipeline that automates software delivery into the classified environment. When something breaks on-site, you fix it. When an update arrives through the data diode or on physical media, you apply it. You are the bridge between xAI's engineering organization and the classified compute facilities where our infrastructure operates.

This role requires significant time on-site at classified compute facilities. You will work closely with customer IT and security teams, cleared facility personnel, and xAI's uncleared platform engineering team (via approved communication channels).

RESPONSIBILITIES:

• Rack, cable, and power GPU server infrastructure (e.g., Dell XE9680L with NVIDIA B200/GB200) and network switching fabric (NVIDIA SN5600, Mellanox QM9700, management switches) in classified data center environments.

• Execute bare metal provisioning using PXE/OSP: deploy squashfs boot images, NVIDIA drivers, MOFED/DOCA packages, and join nodes to Kubernetes clusters (RKE2/kubeadm) — all from pre-staged air-gap bundles with zero internet access.

• Deploy and operate the full Kubernetes platform stack: GPU/network operators, engine-operator, podgroup-operator, xai-scheduler, ingress controllers, storage provisioners, and RBAC.

• Deploy and operate the monitoring and network telemetry stack: VictoriaMetrics, VMAlert, AlertManager, NetCollector, Grafana — configured for local operation without central dependencies.

• Set up and maintain the CDS receive pipeline: data diode receive proxy, local container registry (Harbor), cosign signature verification, and bundle application automation.

• Apply signed software update bundles to classified infrastructure, verify acceptance tests pass, and execute rollback procedures when needed.

• Validate network fabric correctness using LLDP verification, BGP peering checks, and InfiniBand fabric topology validation after initial deployment and hardware changes. Serve as the keyboard operator for network troubleshooting directed by the Network Architect — you execute commands on classified network devices while the architect directs the session on-site or via approved channels.

• Execute compliance and security validation: run STIG scans (OpenSCAP) against deployed systems, verify FIPS 140-3 mode on all nodes, validate AV agent status, and execute pre-admission security checklists before nodes are allowed to serve classified workloads. Document and report compliance status for ATO packages.

• Troubleshoot GPU inference workloads (SGLang, engine-operator, sampling-loadbalancer) in classified environments, working with uncleared engineering teams via approved channels for guidance on complex issues.

• Interface with customer IT, security, and facility teams. Participate in change control board (CCB) processes for classified system modifications. Train customer operations teams on monitoring dashboards, alert response procedures, and basic operational runbooks during deployment handoff.

• Maintain and create operational documentation: site-specific runbooks, deployment validation reports, incident response procedures, and post-deployment handoff materials.

• Participate in on-call rotation for classified site incident response.

• Up to 75% travel to classified compute facilities required.

BASIC QUALIFICATIONS:

• Active Top Secret / SCI (TS/SCI) security clearance with Counterintelligence Polygraph (CI Poly).

• 5+ years of experience in infrastructure engineering, site reliability engineering, or systems engineering, with hands-on datacenter experience (racking, cabling, power, iDRAC/BMC).

• Deep understanding of the Kubernetes stack: container runtimes, CNI (Calico, Cilium), CSI, CRI, Helm, and operator patterns.

• Experience with bare metal Linux provisioning: PXE boot, cloud-init, disk partitioning, driver installation, kernel configuration.

• Proficiency with Infrastructure-as-Code tools (Pulumi, Terraform, or Ansible).

• Experience deploying and operating monitoring stacks (Prometheus, VictoriaMetrics, Grafana, AlertManager).

• Comfortable working independently in classified environments with limited real-time support from uncleared teams.

• Excellent communication and documentation skills — you will be the primary interface between classified operations and uncleared engineering.

PREFERRED SKILLS AND EXPERIENCE:

• Experience with NVIDIA GPU infrastructure: driver installation, CUDA, NCCL, InfiniBand/RoCE, ConnectX NICs, BlueField DPUs.

• Experience with air-gapped or disconnected deployments where all software must be pre-staged.

• Familiarity with network switch configuration and troubleshooting (Cumulus/NVUE, Junos, EOS, NX-OS).

• Experience with cross-domain solutions (CDS), data diodes, or secure transfer mechanisms in classified environments.

• Experience with container image signing and verification (cosign, Sigstore, SBOM tooling).

• Familiarity with RKE2, k3s, or kubeadm for Kubernetes cluster bootstrapping.

• Experience working in SCIF or other classified facility environments.

• Hands-on experience with DISA STIG scanning (OpenSCAP, SCAP Compliance Checker), FIPS 140-3 validation, and CIS benchmark execution (kube-bench).

• Experience with ATO package preparation and change control board (CCB) processes in classified environments.

COMPENSATION AND BENEFITS:

$180,000 - $440,000 USD

Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.

xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Why Apply Through MisuJob?

AI-Powered Job Matching: MisuJob uses advanced artificial intelligence to analyze your skills, experience, and career goals. Our matching algorithm compares your profile against thousands of job requirements to find positions where you have the highest chance of success. This saves you hours of manual job searching and ensures you only see relevant opportunities.

One-Click Applications: Once you create your profile, applying to jobs is effortless. Your resume and cover letter are automatically tailored to highlight the most relevant experience for each position. You can apply to multiple jobs in minutes, not hours.

Career Intelligence: Beyond job matching, MisuJob provides valuable career insights. See how your skills compare to market demands, identify skill gaps to address, and understand salary benchmarks for your experience level. Make data-driven decisions about your career path.

Frequently Asked Questions

How do I apply for this position?

Click the "Register to Apply" button above to create a free MisuJob account. Once registered, you can apply with one click and track your application status in your dashboard.

Is MisuJob free for job seekers?

Yes, MisuJob is completely free for job seekers. Create your profile, get matched with jobs, and apply without any cost. We help you find your dream job without any hidden fees.

How does AI matching work?

Our AI analyzes your resume, skills, and experience to understand your professional profile. It then compares this against job requirements using natural language processing to calculate a match percentage. Higher matches mean better fit for the role.

Can I apply to jobs in other countries?

Absolutely. MisuJob features jobs from companies worldwide, including remote positions. Filter by location or look for remote opportunities to find jobs that match your preferences.

Ready to Apply?

Join thousands of job seekers using MisuJob's AI to find and apply to their dream jobs automatically.

Register to Apply