HIERARCHICAL AUTO EVALUATION OF GENERATIVE AI SYSTEMS
卷宗概要
发明人
Jineet Hiren DOSHI; Maya Vered LIVSHITS; Na XU; Yuan ZHOU; Jeyendran BALAKRISHNAN
IPC 分类
CPC 分类
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.
原文(中文)
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.