Lv62
2064 积分 2024-02-04 加入
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
6天前
已完结
OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration
6天前
已完结
Enhancing Transformer Inference Efficiency on FPGA Through Fully Fusion and Integer-Only Quantization Techniques
21天前
已完结
XShift: FPGA-efficient Binarized LLM with Joint Quantization and Sparsification
1个月前
已完结
Efficient Pruning and Acceleration of Encoder-Based LLM Transformers on eFPGAs
1个月前
已完结
QLlama: An FPGA-Based Microscaling Quantization Accelerator for Energy-Efficient Llama2 Inference
1个月前
已完结
TFLOP: Towards Energy-Efficient LLM Inference An FPGA-Affinity Accelerator with Unified LUT-based OPtimization
1个月前
已完结
dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language Models
1个月前
已完结
MoE-OPU: An FPGA Overlay Processor Leveraging Expert Parallelism for MoE-based Large Language Models
1个月前
已完结
Sparseloop: An Analytical Approach To Sparse Tensor Accelerator Modeling
1个月前
已完结