Vispark logo

Vispark

Vision Small

Canonical ID: vispark/vispark/vision-small

Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.

Reasoning: Yes Tool use: Yes

Overview

text image audio video pdf
Provider
Vispark logo Vispark
Family
vision-small
Release date
2024-05-15
Last updated
Input modalities
text, image, audio, video, pdf
Output modalities
text

Pricing and Limits

Token pricing

Input
$1.05/1M
Output
$3.16/1M
Cached input
N/A
Currency: USD

Context and output

Context window
1,000,000
Max output tokens
65,536
Qualified ID
vispark/vispark/vision-small

Model Summary

Vision Small is a model listing in the Vispark provider catalog. It is categorized under the vision-small family.

The model supports text, image, audio, video, pdf modalities across input/output paths. Reported context capacity is 1,000,000. Pricing is listed at $1.05/1M input and $3.16/1M output.

More From Vispark

Related models from the same provider catalog.

View all from Vispark