Open-weight & owned stack · quantization-efficiency
NVIDIA expanded local AI support to GPUs with at least 24GB of VRAM and highlighted vLLM and llama.cpp optimizations.
NVIDIA expanded local AI support to GPUs with at least 24GB of VRAM, while vLLM and llama.cpp optimizations reportedly improved compute performance by up to 1.9x.
Read the original at wccftech.comOpens the publisher's site in a new tabMore in Open-weight & owned stack
Abliteration.ai released API-accessible GLM-derived models with safety refusals removed.Sep 1MBZUAI’s Institute of Foundation Models released six fully open K2 Horizon models with weights, code, data, and methods.Sep 3Moonshot AI released Kimi K3 open weights, a technical report, and supporting training infrastructure.Sep 2NVIDIA reportedly released Nemotron 3 Ultra, a 550-billion-parameter open-weight model.Sep 3DeepSeek released its V4 vision model and associated open-weight artifacts.Sep 1