Intel's LLM-Scaler project that was born out of their Project Battlematrix initiative aims to make it easier to run generative AI on Arc (Pro) B-Series graphics cards with the likes of vLLM, ComfyUI, SGLang, and other popular AI software in this Docker-based pre-configured AI stack. This week new LLM-Scaler releases brought same-day support for new models and other enhancements...

Intel's LLM-Scaler project that was born out of their Project Battlematrix initiative aims to make it easier to run generative AI on Arc (Pro) B-Series graphics cards with the likes of vLLM, ComfyUI, SGLang, and other popular AI software in this Docker-based pre-configured AI stack. This week new LLM-Scaler releases brought same-day support for new models and other enhancements.
Most notable with the Intel LLM-Scaler-vLLM beta 0.21.0-b3 release on Monday was delivering same-day support for Meta's new Muse Glimmer 30B model. Muse-Glimmer-30B with FP8 online quantization is supported by LLM-Scaler-vLLM on the likes of the Arc Pro B70.
The new LLM-Scaler-vLLM beta also adds suppport for DFlash for Muse-Glimmer-30B and Qwen3.6-27B. There is also better time-to-first-token performance for Gemma-4-31B and Gemma-4-26B-A4B-it. Plus various bug fixes for this updated vLLM stack for Intel graphics. See this GitHub release for those details.
Released today was LLM-Scaler-Omni beta 0.2.0-b1. This new LLM-Scaler-Omni Docker container upgrades to the ComfyUI 0.31 XPU stack, adds support for MiniMax H3 local video generation, supports Wan Animate 2 on the Arc Pro B70 and B60, and expands optimized model coverage with Wan 2.2 14B T2V Turbo, LTX-2, Z-Image / Lumina, and Krea2. This update also adds managed GGUF Q4_1 support and pinned ComfyUI-GGUF-XPU integration.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Lemonade 11.6 Integrates Muse-Glimmer 30B, Experimental TheNoise ROCm Image Generation | 0 | 16.09 | 14-08-2026 |
| 2 | TileRT - Tile-Based Runtime for Ultra-Low-Latency LLM Inference | 0 | 35 | 28-06-2026 |
| 3 | Intel XPU Manager 2.1 Released For Monitoring Arc Pro Graphics On Windows/Linux | 0 | 8.41 | 13-08-2026 |
| 4 | Linux Foundation & Others Launch "Akrites" To Defend Open-Source Software From AI-Enabled Exploits | 0 | 7 | 25-06-2026 |
| 5 | Meta Releases Muse Glimmer, a 30-Billion-Parameter Open-Weight AI Model That Runs on a Single Consumer GPU | 0 | 5.06 | 11-08-2026 |
| 6 | Что такое LiteLLM и для чего его едят | 0 | 6.07 | 29-07-2026 |
| 7 | vLLM vs LMDeploy vs Triton: обзор бэкендов для инференса LLM | 0 | 7 | 18-07-2026 |
| 8 | Hardware-aware framework accelerates large language models without additional training | 0 | 8.57 | 06-08-2026 |
| 9 | Generative AI using Elastic and Amazon SageMaker JumpStart | 0 | 6.25 | 25-07-2023 |