RepoNVIDIANVIDIApublished Jul 7, 2026seen 1w

NVIDIA/nccl-extensions

Cuda

Open original ↗

Captured source

source ↗
published Jul 7, 2026seen 1wcaptured 1whttp 200method plain

NVIDIA/nccl-extensions

Description: Communication patterns for AI, built on top of NCCL device and host APIs

Language: Cuda

License: NOASSERTION

Stars: 0

Forks: 0

Open issues: 0

Created: 2026-07-07T16:06:41Z

Pushed: 2026-07-15T21:25:43Z

Default branch: main

Fork: no

Archived: no

README:

NCCL Extensions

NCCL Extensions is a repository of communication patterns for AI use cases, built on top of NCCL device and host APIs. It speeds up tensor communication for workloads like MoE token shuffle and reinforcement learning weight rollout.

This is an evolving space, and the content here is under constant development and subject to change. We will continue exploring it and welcome your contributions.

What's Inside

[nccl_ep/](nccl_ep/) — Expert Parallelism

Optimized dispatch and combine primitives for Mixture-of-Experts (MoE) token routing, built on NCCL's Device API (LSA and GIN operations).

[nccl_m2n/](nccl_m2n/) — Mesh-to-Mesh Rollout

Experimental library for resharding a tensor between two disjoint groups of GPU processes (e.g. trainer and inference ranks) in a single, zero-copy call, built on NCCL's window API.

Getting Started

This repo vendors NCCL as a git submodule. Clone with:

git clone --recursive

(or git submodule update --init --recursive after a normal clone). See each subproject's README for build instructions.

Contributing

We welcome contributions! See [CONTRIBUTING.md](CONTRIBUTING.md) to get started.

License

This project is licensed under the Apache License, Version 2.0 — see [LICENSE.txt](LICENSE.txt) for details.