image
Bagel
7B-parameter multimodal model unifying text-to-image generation, image editing, and image understanding.
Studio is in early access; availability resolves per project after sign-in.
- Endpoints
- 3
- Eligible
- 1
- Input schemas imported
- 3 / 3
- Delivery
- fal.ai
- Weights
- Not stated
- Developer
- Not stated in source
Overview
Bagel is described as a 7B-parameter multimodal model from ByteDance-Seed that generates both text and images. One API covers text-to-image generation, image-to-image editing, and image understanding including image-to-JSON structured extraction. The unified design serves both text and image tasks together.
Capabilities
- Text-to-image generation
- Image-to-image editing
- Image understanding to structured JSON
- Unified text-and-image API
Best for
- Prompt-to-image creation
- Guided image transforms
- Structured visual extraction
Use cases
- Multimodal content pipelines
- Image-to-JSON data capture
- Edit-plus-understand workflows
Endpoints
2 under review · 1 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.
| Endpoint | Task | Catalog status | Inputs |
|---|---|---|---|
fal-ai/bagel | unknown | Under review | 4 fields · requires prompt |
fal-ai/bagel/edit | image editing | Eligible | 5 fields · requires prompt, image_url |
fal-ai/bagel/understand | unknown | Under review | 3 fields · requires image_url, prompt |
Questions
What tasks does one model cover?
Generation, editing, and understanding across text and images.
What structured output exists?
Image-to-JSON analysis and extraction.