←Back to feed

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Anyone-ai·Argentina·Ashby·REMOTE
13

Department

Software Engineering

Team

Software Engineering

Type

Contract

Posted

Sep 15, 2026

Compensation

$65 per hour

Apply →

Description

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads. We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking. WHAT YOU’LL WORK ON You’ll work with GPU and accelerator kernel tasks involving: - Kernel implementation and debugging - CUDA and Triton optimization - Translation between kernel frameworks - Hardware migration - Operator fusion - Performance profiling and benchmarking - Numerical correctness verification - Compilation and runtime debugging - Memory hierarchy optimization - Kernel-level AI workload performance You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware. WHAT WE’RE LOOKING FOR - 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels - Strong experience with at least two of the following: - CUDA - Triton - NKI / AWS Neuron - Pallas / JAX - Strong understanding of GPU performance optimization - Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers - Understanding of: - Memory bandwidth - Compute throughput - GPU occupancy - Shared memory - Register pressure - Memory coalescing - Bank conflicts - Strong understanding of floating-point numerical correctness and tolerance thresholds - Experience debugging kernel compilation and runtime issues - Ability to distinguish software defects, environment problems, and genuine optimization challenges RELEVANT EXPERIENCE Candidates should have experience with several of the following types of work: - Writing kernels from technical specifications - Translating kernels between CUDA, Triton, or other frameworks - Migrating kernels across hardware platforms - Debugging incorrect kernel implementations - Optimizing kernel performance - Fusing multiple operations into optimized kernels NICE TO HAVE - Experience across both NVIDIA GPU and custom accelerator ecosystems - Experience with AWS Trainium, TPU, JAX, or other accelerators - Compiler engineering experience - Familiarity with MLIR, XLA, or intermediate representation lowering - Contributions to GPU or ML kernel libraries - Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls - Experience with AI model evaluation, RLHF, or technical benchmark development WHAT YOU’LL BE RESPONSIBLE FOR - Reviewing GPU and accelerator kernel implementations for correctness - Comparing outputs against reference implementations - Evaluating numerical tolerance thresholds - Reviewing kernel benchmarks and determining whether comparisons are fair - Identifying performance bottlenecks and optimization opportunities - Assessing whether performance targets are realistic given hardware limits - Reviewing kernel translations and hardware migrations - Identifying compilation, driver, memory, shape, and runtime issues - Determining whether technical tasks are genuinely difficult or incorrectly configured - Providing clear, actionable technical feedback ENGAGEMENT Work Type: Remote Engagement: Part-time, project-based consulting Focus: GPU kernels, performance engineering, debugging, and technical evaluation This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.