HIERARCHICAL AUTO EVALUATION OF GENERATIVE AI SYSTEMS
개요
발명자
Jineet Hiren DOSHI; Maya Vered LIVSHITS; Na XU; Yuan ZHOU; Jeyendran BALAKRISHNAN
IPC 분류
CPC 분류
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.
원문 (중국어)
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.