SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Published in IJCNLP 2025, 2025
Semantic-aware KV cache sharing framework for reducing LLM inference cost across semantically similar prompts.
Recommended citation: Xinye Zhao and Spyridon Mastorakis. (2025). "SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching." arXiv / IJCNLP 2025.
Download Paper
