Open-weight & owned stack · Quantization efficiency
Community Adds Draft-Model Offload For DGX Spark
A community repository enables TCP- and RDMA-based offloading of speculative-decoding draft models from DGX Spark systems to spare GPUs.
Read the original at reddit.comOpens the publisher's site in a new tab