aiDAPTIV TM

Faster Inference and Larger LLM Training, Done Privately On-Prem

Overview

When memory is the limit, compute is not enough

Modern AI often hits a memory wall before it runs out of compute.

Pascari aiDAPTIV™ is a purpose-built Phison solution for AI systems, from client PCs and workstations to edge deployments and servers. It combines aiDAPTIV Middleware and aiDAPTIV Cache Memory to extend usable AI memory across GPU memory, system memory, and a dedicated flash capacity tier.

Run larger models. Retain more context.

AI workloads demand more memory than you think

A model may load cleanly until the real work begins. Add a long document, RAG material, tool definitions, agent state, or a larger MoE model, and the same system can run out of room.

Model weights consume capacity before the first prompt is processed

KV cache grows with conversations, documents, and retrieved context 

Agent workflows carry instructions, tools, observations, and working artifacts across many steps

MoE models need the full expert pool available even though only a subset is active for each token

Fine-tuning adds gradients, optimizer state, activations, and training data

A model may load cleanly until the real work begins. Add a long document, RAG material, tool definitions, agent state, or a larger MoE model, and the same system can run out of room.

A memory system built for AI data

aiDAPTIV Middleware

Helps supported runtimes keep active AI state near compute, retain less-active state in larger memory tiers, and bring it forward when needed again.

aiDAPTIV Cache Memory

The dedicated flash capacity tier for AI state that cannot remain permanently in fast memory without crowding out more immediate work. 

Memory tier
Primary role
Typical data examples 
GPU memory or unified memory
Immediate computation
Active model state, active KV cache, active experts, current training layers
System memory
Staging and expanded capacity on discrete-GPU systems
Prefetched data, intermediate cache, state likely to be needed soon
aiDAPTIV Cache Memory
Retained capacity
Less-active KV cache, colder experts, staged model or training data
Memory Tier

GPU memory or unified memory

Primary role

Immediate computation

Typical data examples 

Active model state, active KV cache, active experts, current training layers

Memory Tier

System memory

Primary role

Staging and expanded capacity on
discrete-GPU systems

Typical data examples 

Prefetched data, intermediate cache, state likely to be needed soon

Memory Tier

aiDAPTIV Cache Memory

Primary role

Retained capacity

Typical data examples 

Less-active KV cache, colder experts, staged model or training data

What aiDAPTIV enables

Keep more context
Run a larger model
Adapt a larger model
What is happening?
Start here

Long prompts, RAG, or repeated documents make time to first token painful 

A sparse MoE model does not fit in available memory

Fine-tuning ends with out-of-memory errors 

My model or total context does not fit, and I am not sure which memory limit is responsible 

Build with aiDAPTIV

Building an AI application?
Building a platform or system?
Looking for technical details?
Looking for a prebuilt workstation?

Explore ABS systems configured with aiDAPTIV on Newegg. 

SEAMLESS INTEGRATION

  • Optimized middleware to extends GPU memory capacity
  • 2x 2TB aiDAPTIVCache to support 70B model
  • Low latency

HIGH ENDURANCE

  • Industry-leading 100 DWPD with 5-year warranty
  • SLC NAND with advanced NAND correction algorithm

aiDAPTIV+ BENEFITS

  • Transparent drop-in
  • No need to change your AI Application
  • Reuse existing HW or add nodes

aiDAPTIV+ MIDDLEWARE

  • Slice model, assign to each GPU
  • Hold pending slices on aiDAPTIVCache
  • Swap pending slices w/ finished slices on GPU

FOR SYSTEM INTEGRATORS

  • Access to ai100E SSD
  • Middleware library license

  • Full Phison support to bring up