Skip to content

EthenEthenEthen

image

Bagel

7B-parameter multimodal model unifying text-to-image generation, image editing, and image understanding.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
3
Eligible
1
Input schemas imported
3 / 3
Delivery
fal.ai
Weights
Not stated
Developer
Not stated in source

Overview

Bagel is described as a 7B-parameter multimodal model from ByteDance-Seed that generates both text and images. One API covers text-to-image generation, image-to-image editing, and image understanding including image-to-JSON structured extraction. The unified design serves both text and image tasks together.

Capabilities

  • Text-to-image generation
  • Image-to-image editing
  • Image understanding to structured JSON
  • Unified text-and-image API

Best for

  • Prompt-to-image creation
  • Guided image transforms
  • Structured visual extraction

Use cases

  • Multimodal content pipelines
  • Image-to-JSON data capture
  • Edit-plus-understand workflows

Endpoints

2 under review · 1 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
fal-ai/bagelunknownUnder review4 fields · requires prompt
fal-ai/bagel/editimage editingEligible5 fields · requires prompt, image_url
fal-ai/bagel/understandunknownUnder review3 fields · requires image_url, prompt

Questions

What tasks does one model cover?

Generation, editing, and understanding across text and images.

What structured output exists?

Image-to-JSON analysis and extraction.