---

title: "ML Systems Performance Engineer (MFU) at Higgsfield AI"

canonical: "https://jobhunter.my/vacancies/job-7d3826a5a8c3aab8-ml-systems-performance-engineer-mfu-higgsfield-ai"

date_posted: "2026-08-24T06:22:01.925Z"

verified: "2026-08-27T01:10:49.805Z"

---

# ML Systems Performance Engineer (MFU)

**Company:** Higgsfield AI

**Location:** Almaty, Kazakhstan

**Remote signal:** Yes

**Published:** 2026-08-24T06:22:01.925Z

**Verified by JobHunter:** 2026-08-27T01:10:49.805Z

## Job description

Why work at Higgsfield AI? Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands. We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like. What you will do • Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration. • Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime. • Optimize distributed training and model-sharding strategies, including data, tensor, pipeline, context, and expert parallelism. • Improve collective communication through topology-aware placement and compute/communication overlap. • Develop or integrate optimized CUDA and Triton kernels • Optimize data loading, preprocessing, sequence packing, and checkpointing so that I/O does not leave accelerators idle. • Diagnose distributed hangs фтв performance regressions. • Improve fault tolerance for long-running training jobs. What we are looking for • Strong experience running and optimizing multi-GPU or multi-node training. • Experience with PyTorch Distributed or an equivalent training framework. • Understanding of GPU architecture, including memory hierarchy, Tensor Cores • Understanding of collective communication, cluster topology, and distributed-training bottlenecks. • Experience with distributed parallelism technologies such as FSDP, DeepSpeed, Megatron-LM, TorchTitan, or similar. • Ability to debug complex performance and reliability problems across multiple layers of the training stack. Nice to have • CUDA, Triton or GPU-kernel development experience. • Experience with NCCL, MPI, UCX, RDMA, InfiniBand, RoCE, GPUDirect, NVLink, or NVSwitch. • Experience training Mixture-of-Experts, multimodal, or reinforcement-learning models. • Knowledge of PyTorch internals, torch.compile, XLA, ML compilers, or custom operators. • Experience with mixed-precision training, including BF16, FP8, or FP4. What We Offer Competitive base salary in USD , based on your experience, skills, and the scope of the role. Equity participation through the company’s stock option program, giving you the opportunity to share in Higgsfield’s long-term growth. Relocation support to Almaty for candidates moving from another city or country. A highly collaborative, fast-paced environment where you can work directly with experienced leaders and have a meaningful impact on the product and company. Opportunities for professional growth, ownership, and career development as the company scales. Company-provided equipment, meals, transportation, or other office benefits. This is a fully on-site role based in our Almaty office . Our team works from the office five days per week for the full working day . We believe in-person collaboration is an important part of how we move quickly, solve complex problems, and build strong teams.

- [Open the employer's vacancy](https://jobs.ashbyhq.com/higgsfieldai/82db7018-fa8a-46fd-a5e1-71c5d9f67fcc/application)

- [Browse the public pilot](https://jobhunter.my/vacancies)

> JobHunter is a monitoring service. Verify availability and application terms on the employer's page.