Close Menu
Alpha Leaders
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
What's On
Why Enterprises Should Consider Open-Weight Models

Why Enterprises Should Consider Open-Weight Models

3 September 2026
AI visionary Ray Kurzweil joins neurotech startup Subsense as product and vision advisor

AI visionary Ray Kurzweil joins neurotech startup Subsense as product and vision advisor

3 September 2026
Before You Buy More Hardware, Check What Your Cluster Is Really Using

Before You Buy More Hardware, Check What Your Cluster Is Really Using

3 September 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Alpha Leaders
newsletter
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
Alpha Leaders
Home » Kioxia AiSAQ Improves AI Inference With Lower DRAM Costs
Innovation

Kioxia AiSAQ Improves AI Inference With Lower DRAM Costs

Press RoomBy Press Room8 July 20253 Mins Read
Facebook Twitter Copy Link Pinterest LinkedIn Tumblr Email WhatsApp
Kioxia AiSAQ Improves AI Inference With Lower DRAM Costs

In April this year, Kioxia’s Rory Bolt gave me a briefing on Kioxia’s AiSAQ, an open-source project intended to promote the expanded use of SSDs in RAG AI solutions. The focus on AI is moving from generating foundational models with massive and expensive training to cost effective and scalable ways to create inference solutions that can solve real world problems.

Retrieval-Augmented Generation is an approach to AI that combined traditional information retrieval systems with large language models. RAG enhances the performance of LLMs by allowing them to access and incorporate information from external knowledge sources, such as databases, websites, and internal documents, before generating a response. This approach helps LLMs produce more accurate, contextually relevant, and up-to-date information, especially when dealing with specific domains or real-time data.

Kioxia has used AI to improve the output of its NAND fabs since 2017, mostly using machine vision to monitor trends and defect rates. In 2020 Kioxia used AI to generate the world’s first AI-designed Manga, Phaedo, drawing on manga drawings and stories based on Osuma Tezuka’s work.

I was told that although larger data centers feed data to their AI models using hard drives, many in-house solutions train using data on SSDs. These solutions often work with foundational LLM models created with very large data sets and use RAG using in-house and perhaps more up to date data to tune the foundational model for a particular application and to avoid hallucinations. The image below illustrates how a database can be used for tuning of the original LLM.

Here the customer query is answered using the LLM as well as domain specific and up to date information in a vector data base. Such RAG solutions can be done with the data base index and vectors all in DRAM, but such an approach can use a lot of memory, making them very expensive, particularly for large data bases.

Microsoft developed Disk ANN which moved the bulk of the vector DB content to SSDs. This reduced the required DRAM footprint for the DB enabling greater scaling of vector DBS. This is used in products such as Azure Vector DB and Cosmos DB.

Kioxia’s All-in-Storage ANNS with Product Quantization, or AiSAQ completes the move of database vectors into storage, further reducing the DRAM requirements. These three approaches are represented in the drawing below.

Kioxia says that this approach enabled greater scalability for RAG workflows and thus better accuracy in the models. The image below shows the significant reduction of DRAM required for large databases compared to the DRAM-based, and DiskANN approach and the improved query accuracy.

In early July Kioxia announced further improvements to its AiSAQ. This new open source release allows flexible controls that allow system architects to define the balance point between search performance and the number of vectors, which are opposing factors with the fixed capacity of SSD storage in the system. The resulting benefit enables architects of RAG systems to fine-tune the optimal balance between specific workloads and their requirements, without any hardware modifications.

Kioxia’s AiSAQ allows more scalable RAG AI inference systems by moving database vectors entirely into storage, thus avoiding DRAM growth with increasing database sizes.

AI AiSAQ Artificial Intelligence DRAM Inference Kioxia NAND RAG Scalability SSD
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link

Related Articles

Why Enterprises Should Consider Open-Weight Models

Why Enterprises Should Consider Open-Weight Models

3 September 2026
Before You Buy More Hardware, Check What Your Cluster Is Really Using

Before You Buy More Hardware, Check What Your Cluster Is Really Using

3 September 2026
The AI Productivity Race Has A False Finish Line

The AI Productivity Race Has A False Finish Line

3 September 2026
The Most Expensive Technology Decisions Are The Ones That Last Years

The Most Expensive Technology Decisions Are The Ones That Last Years

3 September 2026
Build What Better AI Can’t Commoditize

Build What Better AI Can’t Commoditize

3 September 2026
The Promise Of Precision Medicine Is Becoming Practical

The Promise Of Precision Medicine Is Becoming Practical

3 September 2026
Don't Miss
Exclusive: DeFi platform Azura launches after raising .9 million from Initialized

Exclusive: DeFi platform Azura launches after raising $6.9 million from Initialized

By Press Room22 October 2024

Azura, a new platform for decentralized finance, launched on Tuesday after raising $6.9 million in…

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

27 December 2024
Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

22 October 2024
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Latest Articles
The AI Productivity Race Has A False Finish Line

The AI Productivity Race Has A False Finish Line

3 September 20262 Views
AI wants electricity now. The electric grid needs years to catch up

AI wants electricity now. The electric grid needs years to catch up

3 September 20263 Views
The Most Expensive Technology Decisions Are The Ones That Last Years

The Most Expensive Technology Decisions Are The Ones That Last Years

3 September 20262 Views
Meloni breaks Berlusconi’s record for longest-serving uninterrupted government, longest since WWII

Meloni breaks Berlusconi’s record for longest-serving uninterrupted government, longest since WWII

3 September 20262 Views

Recent Posts

  • Why Enterprises Should Consider Open-Weight Models
  • AI visionary Ray Kurzweil joins neurotech startup Subsense as product and vision advisor
  • Before You Buy More Hardware, Check What Your Cluster Is Really Using
  • Canva’s productivity push gains traction—and in Southeast Asia, it’s happening on phones
  • The AI Productivity Race Has A False Finish Line

Recent Comments

No comments to show.
About Us
About Us

Alpha Leaders is your one-stop website for the latest Entrepreneurs and Leaders news and updates, follow us now to get the news that matters to you.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks
Why Enterprises Should Consider Open-Weight Models

Why Enterprises Should Consider Open-Weight Models

3 September 2026
AI visionary Ray Kurzweil joins neurotech startup Subsense as product and vision advisor

AI visionary Ray Kurzweil joins neurotech startup Subsense as product and vision advisor

3 September 2026
Before You Buy More Hardware, Check What Your Cluster Is Really Using

Before You Buy More Hardware, Check What Your Cluster Is Really Using

3 September 2026
Most Popular
Canva’s productivity push gains traction—and in Southeast Asia, it’s happening on phones

Canva’s productivity push gains traction—and in Southeast Asia, it’s happening on phones

3 September 20262 Views
The AI Productivity Race Has A False Finish Line

The AI Productivity Race Has A False Finish Line

3 September 20262 Views
AI wants electricity now. The electric grid needs years to catch up

AI wants electricity now. The electric grid needs years to catch up

3 September 20263 Views

Archives

  • September 2026
  • August 2026
  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • March 2022
  • January 2021
  • March 2020
  • January 2020

Categories

  • Blog
  • Business
  • Entrepreneurs
  • Global
  • Innovation
  • Leadership
  • Living
  • Money & Finance
  • News
  • Press Release
© 2026 Alpha Leaders. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.