((((sandro.net))))
Manuntençao para Pcs
domingo, 2 de agosto de 2026
Show HN: I get 25 deep researched ideas with one single prompt https://ift.tt/GcmQ4AF
Show HN: I get 25 deep researched ideas with one single prompt Make a three-layer workflow with 19 agents work together in parallel with one single prompt, then use that to research anything I want. If I don't specify any topic, it can prompt me some topics and wait for my response. https://ift.tt/0fPtNlZ August 2, 2026 at 12:28AM
Show HN: CostPerPrompt – Live AI API pricing and real-workload cost calculators https://ift.tt/yBjXTOY
Show HN: CostPerPrompt – Live AI API pricing and real-workload cost calculators https://ift.tt/AKDUoQw August 1, 2026 at 10:56PM
sábado, 1 de agosto de 2026
Show HN: Sanitizer – Strip sensitive data from documents locally before an LLM https://ift.tt/vOTRlYE
Show HN: Sanitizer – Strip sensitive data from documents locally before an LLM https://ift.tt/DJsouOV July 31, 2026 at 10:42PM
sexta-feira, 31 de julho de 2026
Show HN: The fastest JSON parser for Go https://ift.tt/hc9I1sk
Show HN: The fastest JSON parser for Go https://ift.tt/QI8Yjqd July 31, 2026 at 12:47AM
quinta-feira, 30 de julho de 2026
Show HN: Edge Drop- #1 productivity and unique clipboard 200 stars on GitHub https://ift.tt/OJE2DGc
Show HN: Edge Drop- #1 productivity and unique clipboard 200 stars on GitHub https://ift.tt/a8cdhNM July 30, 2026 at 03:03AM
Show HN: Alarium – a live 3D flight tracker built on community ADS-B https://ift.tt/Dod8ANs
Show HN: Alarium – a live 3D flight tracker built on community ADS-B https://ift.tt/0VQzyU3 July 30, 2026 at 12:09AM
Show HN: RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop https://ift.tt/QSo6V8r
Show HN: RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop runNburn is an Apache-2.0 Rust inference engine for quantized GGUF models that are too big for your fast memory. The core idea: weights stay file-backed (mmap), host residency stays under an explicit byte budget (--ram-budget), and GPU caches are sized from detected free/total VRAM — never from device-name presets. There is no conversion step, no sidecar cache files, no silent requantization. The GGUF on disk is the single source of truth. The result that made me want to post this: Tencent's Hy3 (295B total / 21B active sparse MoE, a single 97.8 GiB Q2_K GGUF) runs on my desktop with 64 GB of RAM and one consumer NVIDIA GPU. The file is larger than RAM and VRAM combined; the selected experts for each token are pulled on demand (the newest path batches O_DIRECT reads through io_uring), while the pretrained routing is left untouched. On the same machine, same prompt, same decode length, a warm-run median gave ~5.5 tok/s decode vs ~2.0 tok/s for llama.cpp. To be upfront about scope: for models that fit comfortably in VRAM, llama.cpp is still faster than runNburn today — its CUDA kernels have years of tuning and we measure against it honestly (interleaved A/B runs, medians, and any "speedup" that changes output quality is rejected). runNburn's lane is the model that doesn't fit. What's in the box: - CLI, interactive chat, and an OpenAI-compatible server (chat/completions + responses + conversations, SSE streaming, stateful continuation with KV/SSM snapshot reuse). It's built as a single-owner personal server — one active generation is the optimization unit; continuous batching and multi-tenant throughput are explicit non-goals.
- Architecture-aware paths: Llama family, Phi, Gemma, Qwen dense/hybrid/MoE (including GatedDeltaNet layers), Nemotron-H MoE, Hy3, GLM — plus in-model multi-token prediction (self-speculative decoding) with device-side verification where the GGUF ships a drafter.
- Backends: CPU is the default (x86 AVX2, ARM NEON), CUDA and Metal are active, Vulkan/OpenCL are experimental. Android works through a small C ABI (rnb.h).
- Native quantized kernels for Q2_K–Q6_K, Q4_0, Q8_0 — including the low-bit K-quants that big-MoE builds actually ship in. It's pre-1.0 and rough in places; recognition of an architecture doesn't mean every community variant works. But if you've got a model file bigger than your machine and you'd rather it run slowly than not at all, that's exactly the case it was built for. Happy to answer questions about the offloading design, the expert-streaming path, or the measurement protocol. https://ift.tt/vOm2RKY July 29, 2026 at 10:30PM
Assinar:
Postagens (Atom)
DJ Sandro
http://sandroxbox.listen2myradio.com