Pegasus logo
Aman Sachan papers • projects • research index
Papers archive

Papers archive

A complete archive of Aman Sachan’s research papers and manuscript versions.

Browse the archive

Use these pages when you want a cleaner directory view than the homepage.

Paper New LLM Foundry

LLM Foundry: a modular framework for orchestration, memory, compression, tools, and evaluation

A long-form draft for the repository. It formalises the architecture in the repo README, usage guide, and technical paper notes.

draft paper · public demo

adaptersmemoryevaluation
Paper Featured v3 — 32-page manuscript

KVQuant: Adaptive Long-Context KV-Cache Compression for Memory-Bounded LLM Inference

The most complete version: rate-distortion framing, sink-aware allocation, head/layer coupling, benchmark plan, failure analysis, and a full appendixed derivation.

submission-ready preprint

LLM inferencequantisationmemory systems
Paper v2 — extended draft

KVQuant: Adaptive Long-Context KV-Cache Compression for Memory-Bounded LLM Inference

A strong intermediate draft with an expanded mathematical treatment and a more detailed systems analysis.

research manuscript

draftsystemscompression
Paper deep draft

KVQuant: Adaptive Long-Context KV-Cache Compression for Memory-Bounded LLM Inference

An earlier deep write-up, useful for tracing how the mathematics and implementation story evolved.

technical report

technical reportimplementation
Paper concise preprint

KVQuant: Adaptive Long-Context KV-Cache Compression for Memory-Bounded LLM Inference

A shorter, crisper manuscript if you want the idea in a more compact form.

shorter paper version

preprintcompactpaper