Efficient LLMs for Resource-Constrained Edge

Description:

Designing approximate attention mechanisms to cut LLM inference-time compute/memory usage with minimal quality loss, enabling deployment on edge/mobile/low-resource devices.