CNIPA.AI
검색으로 돌아가기
기록

MACHINE-LEARNED ARCHITECTURE FOR STRUCTURED SYNTHETIC DATA GENERATION

발명심사 중
1조회수
20청구항 · 3 독립항
§ Ⅰ

개요

발명자

Aditi Shreya; Manan Dey; Dharani Gopal Akkiraju; Joao Tiago Azevedo Belo; Hariharan Mani

IPC 분류

G6F 9/448G6F 9/445G6N 3/45

CPC 분류

G6F9/4488G6F9/44505G6N3/45

Techniques may generate realistic synthetic data by programmatically generating a configuration file object type and relationship data. This configuration file may be used to retrieve source data matching the object type(s) and/or specific records indicated by the configuration file. The techniques may detect and anonymize private/proprietary information and may determine statistical characteristic(s) of the source data. A batch of prompt(s) may be generated using the source data, the statistical characteristic(s), and the configuration file and may be transmitted to one or more instances of a transformer-based machine-learned model. Sets of synthetic data received from the model instance(s) may be de-duplicated, checked for similarity to the source data (e.g., via embedding the synthetic data and the source data), and may be used to generate synthetic object(s) using the relationship(s) and/or other data indicated by the configuration file. These synthetic object(s) may then be deployed in a software environment.

원문 (중국어)

Techniques may generate realistic synthetic data by programmatically generating a configuration file object type and relationship data. This configuration file may be used to retrieve source data matching the object type(s) and/or specific records indicated by the configuration file. The techniques may detect and anonymize private/proprietary information and may determine statistical characteristic(s) of the source data. A batch of prompt(s) may be generated using the source data, the statistical characteristic(s), and the configuration file and may be transmitted to one or more instances of a transformer-based machine-learned model. Sets of synthetic data received from the model instance(s) may be de-duplicated, checked for similarity to the source data (e.g., via embedding the synthetic data and the source data), and may be used to generate synthetic object(s) using the relationship(s) and/or other data indicated by the configuration file. These synthetic object(s) may then be deployed in a software environment.