Artifacts I have published, with what I contributed, what they derive from, and the evidence behind them. Every entry names its base model and license, because a derivative work is not the same as an original one.
2026-08-11
A 49-billion-parameter model fitted to a single 24 GiB RTX 4090 by measuring per-layer quantization damage and solving the bit allocation against the budget, instead of applying a preset. At the same file size it drifts less from the full-precision original than the standard community quantization, and ties it on five capability benchmarks.
Base model
Relation
quantized
License
NVIDIA Open Model License · Llama 3.3 Community License
My work
The sensitivity map, the mixed-precision recipe, the pack, and the evaluation record. The weights are NVIDIA's. Built with Llama.