CNIPA.AI
검색으로 돌아가기
기록

MACHINE LEARNING CLUSTERING OF EMBEDDINGS CREATED FOR CATEGORICAL DATA USING LARGE LANGUAGE MODELS

발명심사 중
2조회수
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Sumit KUMAR; Prasad MHATRE; Danny BUTVINIK

IPC 분류

G6N 3/91

CPC 분류

G6N3/91

An autonomous machine learning (ML) system and methods are provided that are configured to intelligently cluster categorical data based on embeddings created by prompting a large language model (LLM). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform embedding generation operations which include accessing a data set for categorical data, determining a row of the data set, generating a data container corresponding to the row and an instruction to the LLM that requests an embedding for the row, prompting the LLM to create the embedding using the data container, reducing a dimensionality of the embedding, and outputting the reduced dimensionality embedding to an ML training application executing for training an ML clustering model.

원문 (중국어)

An autonomous machine learning (ML) system and methods are provided that are configured to intelligently cluster categorical data based on embeddings created by prompting a large language model (LLM). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform embedding generation operations which include accessing a data set for categorical data, determining a row of the data set, generating a data container corresponding to the row and an instruction to the LLM that requests an embedding for the row, prompting the LLM to create the embedding using the data container, reducing a dimensionality of the embedding, and outputting the reduced dimensionality embedding to an ML training application executing for training an ML clustering model.