MACHINE LEARNING CLUSTERING OF EMBEDDINGS CREATED FOR CATEGORICAL DATA USING LARGE LANGUAGE MODELS
개요
발명자
Sumit KUMAR; Prasad MHATRE; Danny BUTVINIK
IPC 분류
CPC 분류
An autonomous machine learning (ML) system and methods are provided that are configured to intelligently cluster categorical data based on embeddings created by prompting a large language model (LLM). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform embedding generation operations which include accessing a data set for categorical data, determining a row of the data set, generating a data container corresponding to the row and an instruction to the LLM that requests an embedding for the row, prompting the LLM to create the embedding using the data container, reducing a dimensionality of the embedding, and outputting the reduced dimensionality embedding to an ML training application executing for training an ML clustering model.
원문 (중국어)
An autonomous machine learning (ML) system and methods are provided that are configured to intelligently cluster categorical data based on embeddings created by prompting a large language model (LLM). The system includes a processor and a computer readable medium operably coupled thereto, the computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the processor, to perform embedding generation operations which include accessing a data set for categorical data, determining a row of the data set, generating a data container corresponding to the row and an instruction to the LLM that requests an embedding for the row, prompting the LLM to create the embedding using the data container, reducing a dimensionality of the embedding, and outputting the reduced dimensionality embedding to an ML training application executing for training an ML clustering model.