yassa9

SIDE PROJECTS

pc

Side opensource projects I do in low level programming in C++ and CUDA.

dvlt.cu

dvlt.cu output

  • Suckless and single binary CUDA/C++ inference engine for NVIDIA's DVLT,
  • It reconstructs 3D scenes from sets of images:
    • (depth + rays + camera pose => point cloud).
  • no python, no torch, no framework and zero dependency.

frokenizer

frokenizer benchmark chart

  • Zero allocation, zero dependency and header only C++ BPE tokenizer for Qwen,
  • Uses ahead-of-time DFA compilation to eliminate regex backtracking and heap overhead,
  • Reaching GBs/sec tokenization throughput.

qwen600.cu

qwen600.cu demo

  • +500 stars on github.
  • Static add single batch inference engine for QWEN3-0.6B written in pure CUDA C/C++,
  • Faster than llama.cpp by ~8.5%.