# Software Engineer, GPU Kernels

[Riverai](https://gurify.com/jobs?q=Riverai) · Palo Alto, CA · Posted today

[AI & ML](https://gurify.com/jobs/ai-ml)

[Apply on the original posting → (opens in a new tab)](https://job-boards.greenhouse.io/riverai/jobs/4402432009)

## Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

### Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

### About the Role

We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.

You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.

### What You’ll Do

- Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.

- Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.

- Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.

- Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.

- Profile real workloads and integrate improvements into training and inference runtimes.

- Build reproducible benchmarks that verify correctness, gradients, and performance.

### Skills & Qualifications

### Minimum Qualifications:

- Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.

- Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.

- Strong understanding of GPU architecture, memory hierarchies, and parallel execution.

- Proficiency in C++ and Python.

- Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.

- Strong debugging and profiling skills, with a collaborative approach to engineering.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

- Experience optimizing for NVIDIA Blackwell or Hopper GPUs.

- Work on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.

- Experience implementing backward passes and validating gradients.

- Familiarity with FP8, FP4, and quantized weight layouts.

- Experience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.

- Open-source contributions or a track record of shipping substantial kernel optimizations.

### Logistics & Benefits

- Location: Palo Alto, California.

- Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.

- Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.

- Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.

**Live in Riverai’s hiring system.** Read from the company's own applicant tracking system, not reposted from a job board — so it's a real, open requisition rather than an ad that outlived the role.

We remove it as soon as it disappears at source.

## More jobs like this

- SP [Senior Machine Learning Engineer - Artist-First AI Music Lab](https://gurify.com/job/senior-machine-learning-engineer-artist-first-ai-music-lab-at-spotify-9f7eb6befe29) Spotify · New York, NY · last week
- PA [Software Engineer - Edge AI Systems](https://gurify.com/job/software-engineer-edge-ai-systems-at-palantir-dba3b54e7464) Palantir · Seattle, WA · 7 weeks ago
- MO [Senior Product Manager, AI Platform](https://gurify.com/job/senior-product-manager-ai-platform-at-moloco-d3490c0a1857) Moloco · Menlo Park, United States · last week
- SP [Senior Product Manager, Ultrasound AI - DeepHealth](https://gurify.com/job/senior-product-manager-ultrasound-ai-deephealth-297c3f9390d0) San Diego, United States · last week
- ST [CX Manager- ERP Integration Solutions & AI Operations](https://gurify.com/job/cx-manager-erp-integration-solutions-ai-operations-at-stampli-6324441465f2) Stampli · Austin, TX · last week
- MI [Senior Technical Delivery Manager Data and AI](https://gurify.com/job/senior-technical-delivery-manager-data-and-ai-at-matrix-ifs-852d95bd9d78) Matrix Ifs · United States · last week

```json
{"@context":"https://schema.org/","@type":"JobPosting","title":"Software Engineer, GPU Kernels","description":"\u003Cp\u003EAt River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.\u003C/p\u003E\u003Ch3\u003EWho we are\u003C/h3\u003E\u003Cp\u003EWe are scientists, engineers, and builders from the industry\u0026#39;s top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today\u0026#39;s frontier models.\u003C/p\u003E\u003Ch3\u003EAbout the Role\u003C/h3\u003E\u003Cp\u003EWe are looking for exceptional GPU kernel engineers to build the compute primitives behind River\u2019s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.\u003C/p\u003E\u003Cp\u003EYou will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.\u003C/p\u003E\u003Ch3\u003EWhat You\u2019ll Do\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EBuild fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOptimize memory access, tiling, and synchronization to make efficient use of GPU hardware.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EDevelop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EAccelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EProfile real workloads and integrate improvements into training and inference runtimes.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBuild reproducible benchmarks that verify correctness, gradients, and performance.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003ESkills \u0026amp; Qualifications\u003C/h3\u003E\u003Ch3\u003EMinimum Qualifications:\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003EBachelor\u2019s degree in Computer Science, Computer Engineering, or equivalent practical experience.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong understanding of GPU architecture, memory hierarchies, and parallel execution.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EProficiency in C\u002B\u002B and Python.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong foundations in linear algebra, floating-point arithmetic, and numerical computing.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EStrong debugging and profiling skills, with a collaborative approach to engineering.\u003C/li\u003E\u003C/ul\u003E\u003Cp\u003EPreferred Qualifications: (We encourage you to apply even if you don\u0026#39;t meet all of these)\u003C/p\u003E\u003Cul\u003E\u003Cli\u003EExperience optimizing for NVIDIA Blackwell or Hopper GPUs.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EWork on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience implementing backward passes and validating gradients.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EFamiliarity with FP8, FP4, and quantized weight layouts.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EExperience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EOpen-source contributions or a track record of shipping substantial kernel optimizations.\u003C/li\u003E\u003C/ul\u003E\u003Ch3\u003ELogistics \u0026amp; Benefits\u003C/h3\u003E\u003Cul\u003E\u003Cli\u003ELocation: Palo Alto, California.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003ECompensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EBenefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.\u003C/li\u003E\u003C/ul\u003E\u003Cul\u003E\u003Cli\u003EVisa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.\u003C/li\u003E\u003C/ul\u003E","identifier":{"@type":"PropertyValue","name":"Gurify","value":"software-engineer-gpu-kernels-at-riverai-d38198190289"},"url":"https://gurify.com/job/software-engineer-gpu-kernels-at-riverai-d38198190289","datePosted":"2026-09-12","validThrough":"2026-10-27T23:59:59Z","hiringOrganization":{"@type":"Organization","name":"Riverai","sameAs":"https://job-boards.greenhouse.io/riverai"},"directApply":false,"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressCountry":"US","addressLocality":"Palo Alto"}}}
```

```json
{"@context":"https://schema.org/","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Jobs","item":"https://gurify.com/jobs"},{"@type":"ListItem","position":2,"name":"United States","item":"https://gurify.com/jobs/united-states"},{"@type":"ListItem","position":3,"name":"Software Engineer, GPU Kernels","item":"https://gurify.com/job/software-engineer-gpu-kernels-at-riverai-d38198190289"}]}
```
