US Patent 12,360,962 B1
Semantic Data Determination Using a Large Language Model
Security teams have to analyze data from many third-party sources, and that data rarely arrives with standard field names or complete metadata. This patent describes a system using large language models to automatically determine metadata for fields of non-standardized third-party data.
The LLM describes each field from its name and sample data. A semantic data model framework then refines those descriptions into a finer level of detail, people can validate or override the results, and the descriptions are stored for reuse, so later data can be checked for potential security threats.
Issued July 15, 2025. Assigned to CrowdStrike, Inc. Co-inventors: Arnd Korn, Erdem Torkman, Nikola Milicic and Ritesh Puj.
My contributions
- Co-developed the AI system for automated semantic labeling of third-party security data.
- Designed user override mechanisms for generated field descriptions and semantic models.
- Implemented comprehensive test coverage for automated threat detection workflows.
On USPTO, search for 12360962.