HIERARCHICAL AUTO EVALUATION OF GENERATIVE AI SYSTEMS
案件概要
発明者
Jineet Hiren DOSHI; Maya Vered LIVSHITS; Na XU; Yuan ZHOU; Jeyendran BALAKRISHNAN
IPC分類
CPC分類
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.
原文(中国語)
An auto evaluation system for evaluating large language models (LLMs). The auto evaluation system loads a base auto evaluation class with core functionalities, selects one or more metrics for evaluation, extends the base auto evaluation class to create a child class with additional functionalities tailored to the selected metrics. A judge LLM receives the evaluation prompts from the auto evaluation server for response generation and computes evaluation scores for the test LLM.
外部リソース