Lv2
150 积分 2025-11-09 加入
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
2个月前
已完结
E $^{2}$ LLM: Structure-Guided Efficient Inference for LLMs in Distributed Edge-IoT Environments
3个月前
已完结
Game-Based LLM Inference Task Offloading for Edge Computing System
3个月前
已完结
Reliable and Efficient LLM Inference on Resource-Constrained Mobile Devices via Dynamic Scheduling
3个月前
已完结
TightLLM: Maximizing Throughput for LLM Inference via Adaptive Offloading Policy
3个月前
已完结
Research on lightweight optimization of large language models for resource-constrained environments
7个月前
已完结
Edge-LLM: A Collaborative Framework for Large Language Model Serving in Edge Computing
7个月前
已完结
Enhancing LLM QoS through Cloud-Edge Collaboration: A Diffusion-based Multi-Agent Reinforcement Learning Approach
7个月前
已完结
Joint Inference Offloading and Model Caching for Small and Large Language Model Collaboration
8个月前
已完结
Large Language Models (LLMs) Inference Offloading and Resource Allocation in Cloud-Edge Computing: An Active Inference Approach
8个月前
已完结