Needle 2:6.6k stars 的 14MB 设备端 LLM,把 Tool Calling 塞进 28MB 内存,让手机 / 智能家居 / 机器人真正能本地对话
cactus-compute/needle 是 GitHub 6.6k stars、435 forks、Cactus Compute 开源的 14MB 设备端基础模型 Needle 2(45M 参数、CQ2-bit 量化、Simple Attention Network 架构),自带字节级 grammar 约束解码、confidence gating、tool retrieval 与 256-token 滑动窗口,把 Tool Calling / Structured Extraction 跑在 28MB 常驻内存内,支持 LoRA 微调与 .cact 导出,配合本地 Playground 可在手机、嵌入式、智能家居、机器人上离线运行。