logo
Aperçu Nouvelles

nouvelles de l'entreprise Scality's AI Inferencing Factory accelerates inferencing

Certificat
Chine Beijing Qianxing Jietong Technology Co., Ltd. certifications
Chine Beijing Qianxing Jietong Technology Co., Ltd. certifications
Examens de client
Le personnel de vente de Beijing Qianxing Jietong Technology Co.,Ltd sont très professionnel et patient. Ils peuvent fournir des citations rapidement. La qualité et l'emballage des produits sont également très bons. Notre coopération est très lisse.

—— LLC de》 de Festfing DV de 《

Quand je recherchais l'unité centrale de traitement d'Intel et le disque transistorisé de Toshiba instamment, Sandy de Beijing Qianxing Jietong Technology Co., Ltd m'a donné beaucoup d'aide et m'a obtenu les produits que j'ai eus besoin rapidement. Je l'apprécie vraiment.

—— Kitty Yen

Sandy de Beijing Qianxing Jietong Technology Co.,Ltd est un vendeur très soigneux, qui peut me rappeler des erreurs de configuration à temps où j'achète un serveur. Les ingénieurs sont également très professionnels et peuvent rapidement compléter le processus de essai.

—— Strelkin Mikhail Vladimirovich

Nous sommes très satisfaits de notre expérience de travail avec Beijing Qianxing Jietong. La qualité du produit est excellente et la livraison est toujours à l'heure. Leur équipe de vente est professionnelle, patiente et très serviable pour toutes nos questions. Nous apprécions vraiment leur soutien et nous nous réjouissons d'un partenariat à long terme. Fortement recommandé !

—— Ahmad Navid

Qualité: “Grande expérience avec mon fournisseur. Le MikroTik RB3011 était déjà utilisé, mais il était en très bon état et tout fonctionnait parfaitement.et toutes mes préoccupations ont été traitées rapidementUn fournisseur très fiable, très recommandé.

—— Geran Colesio

Je suis en ligne une discussion en ligne
Société Nouvelles
Scality's AI Inferencing Factory accelerates inferencing


Scality’s Maestro manages Artesca fleets. Its AI Inference Factory delivers a shared KV cache pool across on-premises GPU servers, matching Cloudian and MinIO while adding prefill/decode separation to boost GPU efficiency.


This AI Inference Factory combines validated open-weight models, a disaggregated inference-serving layer, and Scality AI Data Infrastructure (ADI). ADI uses policy-governed autonomous operations to manage large datasets across the full AI lifecycle. It gives enterprises, government agencies and neo-cloud providers a supported alternative to cloud AI services, removing the burden of assembling, integrating and maintaining the full software stack. The platform places ADI object storage under a vLLM/Dynamo serving layer as shared KV cache, with S3-over-RDMA and a native KV connector.


dernières nouvelles de l'entreprise Scality's AI Inferencing Factory accelerates inferencing  0

Scality co-founder and CEO Jérôme Lecat said: “With AI moving into mission-critical production environments, organisations need greater control over where inference runs, how their models are managed and what happens to their data. For 15 years, Scality has built data infrastructure relied on by thousands of global customers for 24/7 operation. AI Inference Factory brings that experience to on-premises AI, giving organisations the reliability and sovereignty they need to run critical AI workloads on their own terms.”


The firm notes on-prem inference offers data sovereignty and potential cost benefits. Cloud inference uses per-token pricing, so costs rise as AI workflows grow more useful. Cloud providers may also alter model versions and quantisation at their discretion, disrupting dependent workflows.


Scality’s AI Inference Factory supports specialised chatbots, AI-assisted software development, agentic applications and other inference-heavy workloads. Its stack includes four software layers:
-Validated open-weight models maintained within the supported stack;
-Disaggregated inference serving separating prefill from decode so each can scale independently;
-A control plane that authenticates and meters requests, routes traffic to suitable GPUs holding relevant context and schedules workloads against SLA targets;
-Scality ADI, delivering high-performance object storage for models, enterprise data and inference state.


Scality launched Autonomous Data Infrastructure (ADI) in May as an enterprise-focused data infrastructure management platform. ADI uses policy-driven AI agent workers to place data across four storage tiers defined by performance, cost and protection.


ADI integrates with standard AI stacks and deploys at multi-petabyte scale on NVMe, with RDMA access to HDDs for nearly unlimited KV cache capacity. A single namespace automatically tiers data across TLC, HDD and optional tape, helping organisations match storage performance and cost to each stage of the AI data lifecycle.


ADI acts as the shared storage layer for AI Inference Factory, extending KV cache beyond GPU HBM limits. This lets AI inference infrastructure retain and retrieve model context rather than forcing GPUs to recompute it, lifting GPU utilisation and cutting inference costs.

Scality states ADI delivers a shared multi-petabyte cache readable by GPUs at latency within the same order of magnitude as GPU memory. Though slower, this allows GPUs to fetch existing context and focus on inference. This mechanism is similar to other KV caching systems.


dernières nouvelles de l'entreprise Scality's AI Inferencing Factory accelerates inferencing  1


Co-founder and CTO Giorgio Regni said: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with GPUs staying fully occupied. Storage no longer forces KV cache to reside inside the GPU server.”


He added: “That enables disaggregated serving. One GPU pool handles prefill, processing prompts and writing KV cache to ADI. A second pool handles decode, reading cache back and generating tokens. With Scality ADI as shared cache, any decode GPU can pick up any context, with virtually no size limit. Prefill no longer interrupts decode, and both pools run at full load.”


As Regni describes, the AIIF architecture splits AI inference prefill and decode, using ADI as shared cache between GPU pools. Prefill computes KV cache for new prompts, while decode generates answers token by token. When both run on the same GPUs, new prefill jobs disrupt active decodes. Once separated, each pool handles one task full-time and scales independently, drawing from the same shared KV cache. Available decode GPUs can access stored context without binding it to a specific GPU server. Scality cites two examples showing performance gains from this separation:
DistServe (OSDI 2024) achieved up to 7.4 times more requests at identical latency by moving cache directly between GPUs. In Scality’s design, prefill writes KV cache to ADI and decode reads it back, so any decode GPU can access stored context, with tiered storage imposing almost no cache size limits.
Moonshot AI's Mooncake delivered 75 percent more requests on Kimi’s production traffic by deploying a shared KV cache pool.

Scality AIIF prefill-decode split.


Scality testing shows ADI can reduce GPU consumption while preserving performance:
1.9-second load time for Gemma-3 27B, nearly 10x faster than local NVMe, with the model streamed in parallel across the cluster over RDMA;
166 ms warm time-to-first-token restoring a 14K-token context from ADI, only 83 ms behind HBM;
14x faster KV cache retrieval than recomputation for a 14K-token context, up to 72x faster for a 439K-token context;
KV cache more than 80x larger than one GPU’s memory, supporting up to 1,000 resumable concurrent sessions without recomputation;
No measurable impact on token generation since context is restored before the first token;
97 percent of network line rate for transfers between GPUs and storage.


Scality validates and maintains the full stack. It works with open-source harnesses and agentic frameworks including OpenCode, Hermes, Goose, LangGraph and Pydantic AI, and runs validated open-weight models such as Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek. The software deploys on standard servers from Dell, HPE, Lenovo and Supermicro.


Stack components use open code, letting customers inspect how inference state is stored and moved and submit contributions for Scality review.


Scality AI Inference Factory's single end point.
Scality AI Inference Factory packages the whole serving stack behind one endpoint, including vLLM, Dynamo, the Scality AI Connector, Scality ADI and more. It is available now as a software license or fully managed service.


dernières nouvelles de l'entreprise Scality's AI Inferencing Factory accelerates inferencing  2

                                                                           Nvidia KV caching


Nearly all filesystem storage vendors support Nvidia's STX scheme. It defines hardware requirements including BlueField-4 DPUs, Spectrum-X networking and dedicated flash storage tiers. Vendors in this group include DDN, Everpure, HPE, Hitachi Vantara, IBM, NetApp, Nutanix, VAST Data and WEKA. Object storage (S3) Nvidia CMX partners include Cloudian and MinIO with its petabyte-scale caching system.


Cloudian’s AIDP is a turnkey, on-premises appliance as an alternative to public cloud S3 data sources. This S3-compatible, RDMA-based datastore runs AI models and agents on Nvidia Blackwell GPU hardware and software plus BlueField DPUs, complying with Nvidia’s AI Data Platform reference architecture. It lets customers run production AI on their own IT infrastructure and retain direct control over sensitive data. Cloudian claims up to 60 percent cost savings by avoiding recurring token, egress and inference fees that make large-scale public cloud AI costly.


This value proposition resembles Scality’s, but lacks prefill-decode separation.
MinIO states MemKV lets an entire GPU cluster access a shared context pool at microsecond latencies matching inference speeds, rather than waiting for millisecond-scale latency. Again, this message aligns with Scality’s, without prefill-decode separation.


Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!

Temps de bar : 2026-10-10 10:03:29 >> Liste de nouvelles
Coordonnées
Beijing Qianxing Jietong Technology Co., Ltd.

Personne à contacter: Ms. Sandy Yang

Téléphone: 13426366826

Envoyez votre demande directement à nous (0 / 3000)