flashattention

Star

Here are 21 public repositories matching this topic...

egaoharu-kensei / flash-attention-triton

Star

Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode

Updated Jan 12, 2026
Python

MaxLSB / flash-attn2

Star

FlashAttention for sliding window attention in Triton (fwd + bwd pass)

python deep-learning pytorch triton sliding-window flash-attention-2 flashattention

Updated Jun 25, 2025
Python

lyj20071013 / Triton-FlashAttention

Star

This repository contains multiple implementations of Flash Attention optimized with Triton kernels, showcasing progressive performance improvements through hardware-aware optimizations. The implementations range from basic block-wise processing to advanced techniques like FP8 quantization and prefetching

pytorch triton attention flashattention

Updated Mar 26, 2026
Python

llcuda / llcuda

Star

CUDA 12-first backend inference for Unsloth on Kaggle — Optimized for small GGUF models (1B-5B) on dual Tesla T4 GPUs (15GB each, SM 7.5)

python machine-learning ai deep-learning jupyter gpu cuda inference pytorch nvidia cuda-kernels google-colab tensor-cores tesla-t4 llm gguf unsloth flashattention

Updated Feb 1, 2026
Jupyter Notebook

Any-Winter-4079 / Nano-GPT-Speedrun-Track

Star

This repo represents my Nano-GPT speedrun playground, which started coding along Let's reproduce GPT-2 (124M), then moved into further improvements.

decoder transformers speedrun muon rope gpt-2 gpt3 decoder-model nanogpt flexattention flashattention

Updated Feb 28, 2026
Python

Wulfic / AI-OS

Star

HRM-sMoE LLM training toolkit.

Updated Dec 18, 2025
Python

XiaomingFun233 / flash_attn_cuda

Star

easy naive flash attention without optimization base on origin paper

decode attention cuda-kernels flashattention

Updated Nov 14, 2025
Cuda

kennedy-kitoko / yolov12-sdpa-flashattention-pytorch

Star

PyTorch implementation of YOLOv12 with Scaled Dot-Product Attention (SDPA) optimized by FlashAttention for fast and efficient object detection.

pytorch yolo object-detection sdpa ultralytics yolov12 flashattention

Updated Jun 20, 2025
HTML

aidendorian / Marcella-60M-SLM

Star

A 66M parameter decoder-only transformer language model implemented from scratch in PyTorch. Features a custom SentencePiece tokenizer, RoPE positional embeddings, SwiGLU feed-forward network, per-layer KV cache for efficient autoregressive inference, and a Svelte-based streaming chat interface.

transformers torch pytorch pretrained-models slm language-model rope alpaca sdpa finetuning sentencepiece transformer-models kv-cache small-language-models fineweb flashattention marcella

Updated Apr 5, 2026
Python

rogerchang1108 / FlashAttention-with-CUDA

Star

200 lines Flash Attention (only forward pass) in CUDA.

cuda forward-pass flashattention

Updated Feb 23, 2025
Cuda

kalyani-25 / Reimplementation_flash-attention-from-scratch

Star

16-step CUDA optimization of FlashAttention-2 achieving 99.2% of official performance on A100 — Ampere architecture

deep-learning cuda pytorch ampere gpu-kernels nsight llm-inference flashattention

Updated Mar 6, 2026
Cuda

LessUp / llm-speed

Star

CUDA kernels for LLM inference: FlashAttention forward, Tensor Core GEMM, PyTorch bindings, and benchmarkable reference implementations.

benchmarking cuda gpu-acceleration attention cuda-kernels gemm pytorch-extension tensor-core llm-inference flashattention

Updated Apr 23, 2026
Python

MayurVijayPatil / amd-llm-rocm

Star

White paper & reproducible benchmark suite for LLM inference optimization on AMD MI300X using ROCm 6.1

benchmark amd hip quantization rocm awq vllm llm-inference mi300x flashattention

Updated Apr 17, 2026
Jupyter Notebook

LessUp / hpc-ai-optimization-lab

Star

CUDA kernel optimization lab: GEMM, FlashAttention, quantization, and GPU performance learning.

cuda high-performance-computing cuda-kernels gpu-computing gemm cpp20 gpu-programming ai-inference tensor-core nanobind flashattention kernel-optimization

Updated Apr 27, 2026
Cuda

puneethkotha / vit-inference-optimization

Star

ViT-L/16 inference optimization - 4-bit NF4 quantization, FlashAttention-2 vs SDPA benchmarking, 40.5% latency reduction

python optimization cuda transformers inference pytorch vit quantization inference-optimization vision-transformer flashattention

Updated Mar 17, 2026
Jupyter Notebook

tmjoshi / triton_FA

Star

FlashAttention2 Analysis in Triton

cuda-kernels triton-kernels flashattention

Updated Oct 24, 2025
Python

Kaminyou / Flash-Attention-Practice

Star

An minimal CUDA implementation of FlashAttention v1 and v2

deep-learning cuda-programming flashattention

Updated Jun 5, 2025
Python

ericithomas / inference_engine

Star

From-scratch CUDA implementation of memory-efficient transformer attention with up to 9.5x speedup over a naive baseline, deployed end-to-end to Raspberry Pi 4.

raspberry-pi time-series cuda pytorch transformer attention-mechanism edge-computing patchtst flashattention

Updated Apr 25, 2026
Jupyter Notebook

adityakamat24 / triton-fast-mha

Sponsor

Star

A high-performance kernel implementation of multi-head attention using Triton. Focused on minimizing memory overhead and maximizing throughput for large-scale transformer layers. Includes clean-tensor layouts, head-grouping optimisations, and ready-to-benchmark code you can plug into custom models.

transformers parallelism triton memory-efficiency gpu-optimization multi-head-attention kernel-programming flashattention

Updated Aug 12, 2025
Python

lyj20071013 / Triton-Ring-FlashAttn

Star

A minimal, educational implementation of Ring Attention logic using custom OpenAI Triton kernels. Supports blockwise computation and online softmax merging.

cuda pytorch triton llm flashattention

Updated Jan 23, 2026
Python

Improve this page

Add a description, image, and links to the flashattention topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the flashattention topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

flashattention

Here are 21 public repositories matching this topic...

egaoharu-kensei / flash-attention-triton

MaxLSB / flash-attn2

lyj20071013 / Triton-FlashAttention

llcuda / llcuda

Any-Winter-4079 / Nano-GPT-Speedrun-Track

Wulfic / AI-OS

XiaomingFun233 / flash_attn_cuda

kennedy-kitoko / yolov12-sdpa-flashattention-pytorch

aidendorian / Marcella-60M-SLM

rogerchang1108 / FlashAttention-with-CUDA

kalyani-25 / Reimplementation_flash-attention-from-scratch

LessUp / llm-speed

MayurVijayPatil / amd-llm-rocm

LessUp / hpc-ai-optimization-lab

puneethkotha / vit-inference-optimization

tmjoshi / triton_FA

Kaminyou / Flash-Attention-Practice

ericithomas / inference_engine

adityakamat24 / triton-fast-mha

lyj20071013 / Triton-Ring-FlashAttn

Improve this page

Add this topic to your repo