{"avoid_when":"笔记本一键试用（更看 Ollama）","content_hash":"031723cf7333b8fceaae364619441f84c3d765703731af386c7a31c65b82c09f","domain":"ai-agents","evidence_urls":["https://github.com/vllm-project/vllm"],"id":"vllm","language":"Python","license":"","name":"vLLM","niche":"high-throughput inference","repo":"https://github.com/vllm-project/vllm","source_updated_at":null,"status":"active","summary":"高吞吐 LLM 推理引擎，PagedAttention，适合自托管服务。","tags":["llm-api","local-llm"],"use_when":"GPU 集群/服务器上要高并发开源模型推理","verification_status":"unverified","verified_at":null}
