Alberto.Codes

HomeAboutExperienceProjectsPublicationsBlogContact

Publications

Artifacts I have published, with what I contributed, what they derive from, and the evidence behind them. Every entry names its base model and license, because a derivative work is not the same as an original one.

Quantized model

2026-08-11

Llama-3_3-Nemotron-Super-49B-v1_5-fit24gib-GGUF

A 49-billion-parameter model fitted to a single 24 GiB RTX 4090 by measuring per-layer quantization damage and solving the bit allocation against the budget, instead of applying a preset. At the same file size it drifts less from the full-precision original than the standard community quantization, and ties it on five capability benchmarks.

Base model

nvidia/Llama-3_3-Nemotron-Super-49B-v1_5

Relation

quantized

License

NVIDIA Open Model License · Llama 3.3 Community License

My work

The sensitivity map, the mixed-precision recipe, the pack, and the evaluation record. The weights are NVIDIA's. Built with Llama.

© 2026 Alberto Nieto. All rights reserved.