Dollarchip Technology Inc.

產品

XRM-SSD V24.5 Gene Fusion Powered Adaptive AI Compute Platform

Call

Real-world benchmark data for XRM-SSD V24.5 Running Llama-2-70B on a single A100 80GB card, V24.5 achieved a peak throughput of 260.1 tokens/second in Turbo-Burst Mode (+69% vs V24.4 baseline).

This performance gain comes from the Gene Fusion engine, which leverages task similarity to dynamically fuse memory hierarchy, scheduling, and instruction paths.
Under high-concurrency conditions, the system can switch to lower-fidelity (χ=64) paths within microseconds, reducing memory traffic by up to 89%.

In overload scenarios, the Crash Snapshot Engine enables state recovery in 0.034 milliseconds.
In separate vLLM-based production-style load tests (100 concurrent requests, max_output_len=50), the system maintained a 100% success rate with a total throughput of 2573 tok/s at 100 concurrency.

Note: The 260.1 tok/s represents the Turbo-Burst Mode ceiling (not steady-state). Sustained throughput under continuous multi-tenant load is approximately 167 tok/s.