Getting Started
vlm-1 extracts structured markdown from long documents and reports.
1
Upload Document
Upload the document with
/v1/files.from vlmrun.client import VLMRun
from vlmrun.client.types import FileResponse
# Initialize the client
client = VLMRun(api_key="<your-api-key>")
# Upload the file
response: FileResponse = client.files.upload(
file=Path("<path/to/test.pdf>")
)
print(f"Uploaded file:\n {response.model_dump()}")
import { VlmRun } from "vlmrun";
// Initialize the client
const client = new VlmRun({
apiKey: "<your-api-key>",
});
// Upload the file
const fileResponse = await client.files.upload(
filePath: "<path/to/test.pdf>"
);
console.log(fileResponse);
Uploaded file:
{
'id': '1e76cfd9-ba99-49b2-a8fe-2c8efaad2649',
'filename': 'file-20240815-7UvOUQ-earnings_single_table.pdf',
'bytes': 62430,
'purpose': 'assistants',
'created_at': '2024-08-15T02:22:06.716130',
'object': 'file'
}
2
Submit the Document AI Job
Submit the uploaded
file_id to /v1/document/generate. For long documents, set batch=True to queue the job.from vlmrun.client.types import PredictionResponse
# Submit the document for parsing
# Note: In this case, we are using the `document.markdown` domain
# which is optimized for extracting structured markdown from documents
response: PredictionResponse = client.document.generate(
file=response.id,
domain="document.markdown",
batch=True,
)
print(f"Document submitted [id={response.id}]")
print(response.model_dump_json(indent=2))
const response = await client.document.generate({
fileId: fileResponse.id,
domain: "document.markdown",
batch: true,
});
console.log(response);
Document parsing job submitted:
{
"id": "052cf2a8-2b84-45f5-a385-ccac2aae13bb",
"created_at": "2024-08-15T02:22:09.157788",
"response": null,
"status": "pending"
}
3
Fetch the Results
Fetch results with
/v1/predictions/{request_id}. The extraction is JSON in response.# Fetch the results
response: PredictionResponse = client.predictions.wait(request_id)
print(f"Document parsing job results:\n {response.model_dump()}")
// Fetch the results
const response = await client.predictions.wait(requestId);
console.log(response);
{
"id": "052cf2a8-2b84-45f5-a385-ccac2aae13bb",
"created_at": "2024-08-15T02:22:09.157788",
"status": "completed",
"response": {
"pages": [
{ // page 0
"content": "<Figure id=\"fg-0\"/>\n\n# Fine-tuning\nTechnique\n\n---\n\nFebruary 2024",
"markdown_content": "<Figure id=\"fg-0\"/>\n\nOpenAI logo\n\n# Fine-tuning\nTechnique\n\n---\n\nFebruary 2024",
"tables": null,
"figures": [
{
"id": 0,
"title": null,
"caption": null,
"content": "OpenAI logo"
}
]
},
{ // page 1
"content": "# Overview\n\nFine-tuning involves adjusting the parameters of pre-trained models on a specific dataset or task. This process enhances the model's ability to generate more accurate and relevant responses for the given context by adapting it to the nuances and specific requirements of the task at hand.\n\n**Example use cases**\n- Generate output in a consistent format\n- Process input by following specific instructions\n\n## What we'll cover\n\n* When to fine-tune\n* Preparing the dataset\n* Best practices\n* Hyperparameters\n* Fine-tuning advances\n* Resources\n\n---\n3",
"markdown_content": "# Overview\n\nFine-tuning involves adjusting the parameters of pre-trained models on a specific dataset or task. This process enhances the model's ability to generate more accurate and relevant responses for the given context by adapting it to the nuances and specific requirements of the task at hand.\n\n**Example use cases**\n- Generate output in a consistent format\n- Process input by following specific instructions\n\n## What we'll cover\n\n* When to fine-tune\n* Preparing the dataset\n* Best practices\n* Hyperparameters\n* Fine-tuning advances\n* Resources\n\n---\n3",
"tables": null,
"figures": null
},
...
]
}
}
See the
MarkdownPage Schema for the document.markdown domain.