snowflake.snowpark.functions.ai_extract¶
- snowflake.snowpark.functions.ai_extract(input: Union[Column, str], response_format: Union[dict, list], scores: Optional[bool] = None, config: Optional[dict] = None) Column[source]¶
Extracts information from an input string or file based on the specified response format.
- Parameters:
input – Either: - A string or Column containing text to extract information from - A FILE type Column representing a document to extract from (e.g. from
to_file())response_format –
Information to be extracted in one of the following formats:
Simple object schema (dict) mapping feature names to extraction prompts:
{'name': 'What is the last name of the employee?', 'address': 'What is the address of the employee?'}Array of strings containing the information to be extracted:
['What is the last name of the employee?', 'What is the address of the employee?']Array of arrays containing two strings (feature name and extraction prompt):
[['name', 'What is the last name of the employee?'], ['address', 'What is the address of the employee?']]Array of strings with colon-separated feature names and extraction prompts:
['name: What is the last name of the employee?', 'address: What is the address of the employee?']JSON Schema format: a dict with a top-level
schemakey whose value is a JSON Schema object ('type': 'object'). Each entry underpropertiesdescribes one field to extract, wheredescriptionis the extraction prompt andtypeis one of'string'(a single value),'array'(a list), or'object'(a table, which must also specifycolumn_orderingand its ownproperties), e.g.:You can’t combine the JSON Schema format with the other formats above: if
response_formatcontains aschemakey, every field to extract must be defined within that JSON Schema.
scores – Optional boolean. When
True, the returned JSON includes ascoringobject alongsideresponse. Each extracted field gets ascorebetween 0 and 1 indicating the model’s confidence in the extracted value. Requires the named-argument SQL syntax; providing this parameter automatically switches to named-argument form. Default isNone(no scores).config –
Optional dict of configuration settings for file inputs. Supported keys:
scale_factor: A numeric value from 1.0 through 4.0 that scales pages before processing, which can enhance OCR quality for dense layouts or small text.
Only valid when
inputis a FILE column. Requires the named-argument SQL syntax; providing this parameter automatically switches to named-argument form.
- Returns:
A Column containing a JSON object with the extracted information. When
scores=True, the object also includes ascoringsub-object with per-field confidence scores.
Note
You can either ask questions in natural language or describe information to be extracted (e.g., ‘City, street, ZIP’ instead of ‘What is the address?’)
To extract a list, add ‘List:’ at the beginning of each question
Maximum of 100 features can be extracted
Documents must be no more than 125 pages long
Maximum output length is 512 tokens per question
Supported file formats: PDF, PNG, PPTX, EML, DOC, DOCX, JPEG, JPG, HTM, HTML, TEXT, TXT, TIF, TIFF, BMP, GIF, WEBP, MD Files must be less than 100 MB in size.
Examples: