Paper page - ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

ConsumerBench evaluates GenAI system efficiency and response time on end-user devices through a comprehensive benchmarking framework, emphasizing realistic multi-application scenarios and customizable workflows.

The recent shift in Generative AI (GenAI) applications from cloud-only
environments to end-user devices introduces new challenges in resource
management, system efficiency, and user experience. This paper presents
ConsumerBench, a comprehensive benchmarking framework designed to evaluate the
system efficiency and response time of GenAI models running on end-user
devices. Unlike existing benchmarks that assume exclusive model access on
dedicated GPUs, ConsumerBench simulates realistic multi-application scenarios
executing concurrently on constrained hardware. Furthermore, ConsumerBench
supports customizable workflows that simulate complex tasks requiring
coordination among multiple applications. ConsumerBench captures both
application-level metrics, including latency and Service Level Objective (SLO)
attainment, and system-level metrics like CPU/GPU utilization and memory
bandwidth. Through extensive experiments, ConsumerBench reveals inefficiencies
in resource sharing, unfair scheduling under greedy allocation, and performance
pitfalls of static model server configurations. The paper also provides
practical insights for model developers and system designers, highlighting the
benefits of custom kernels tailored to consumer-grade GPU architectures and the
value of implementing SLO-aware scheduling strategies.

Source link

What's Hot

MIT cinches deal with lender to borrow up to $500 million

Emergence AI’s CRAFT arrives to make it easy for enterprises to automate their entire data pipeline

In just 4 months, AI medical scribe Abridge doubles valuation to $5.3B

Paper page – ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

Paper page – LettinGo: Explore User Profile Generation for Recommendation System

Paper page – Enhancing Step-by-Step and Verifiable Medical Reasoning in MLLMs

Paper page – CommVQ: Commutative Vector Quantization for KV Cache Compression

Ezrom Legae And Art Under Apartheid At High Museum Of Art In Atlanta

Chanel Launches Arts & Culture Magazine

Publicity Wizard Jalila Singerff On The Vital PR Rules For 2025

Tourist Damaged 17th-Century Portrait at Florence’s Uffizi Galleries

MIT cinches deal with lender to borrow up to $500 million

Emergence AI’s CRAFT arrives to make it easy for enterprises to automate their entire data pipeline

In just 4 months, AI medical scribe Abridge doubles valuation to $5.3B

What's Hot

Paper page – ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

Related Posts

Subscribe to Updates