Florence 2

Use Microsoft's Florence 2 multimodal model through our Serverless Cloud API

We support Microsoft's Florence 2, a multimodal vision-language model, via our Serverless Cloud API. Florence 2 supports captioning, object detection, segmentation, and OCR through task prompts (such as <CAPTION>, <OD>, <OCR>, <REFERRING_EXPRESSION_SEGMENTATION>).

Florence 2 pretrained aliases

Use the alias as the model_id in your request and the runtime resolves it to the corresponding pretrained weights.

Alias
florence-2-base
florence-2-large

Florence 2 accuracy

Headline zero-shot metrics from the official model cards (base, large):

Benchmarkflorence-2-baseflorence-2-large
COCO Caption (CIDEr)133.0135.6
COCO detection (mAP)34.737.5
RefCOCO (accuracy)53.956.3

Florence 2 inference speed

Latency measured with Roboflow Inference on 1x NVIDIA L4, batch size 1, generating exactly 128 tokens with greedy decoding from the <CAPTION> prompt. Latency scales with output length, so use tokens/sec to estimate other lengths.

AliasLatency, 128 tokens (ms)Tokens/sec
florence-2-base652198
florence-2-large1120115

Florence 2 API

Florence 2 runs through the shared /infer/lmm endpoint. Call it through the HTTP endpoint directly with curl, or with the inference-sdk wrapper.

1

Get your API Key

Create a Roboflow account, find your key on the Roboflow API settings page and make it available to your shell:

export ROBOFLOW_API_KEY="your-key-here"
2

Run the model

Call the /infer/lmm endpoint with a task prompt using curl:

curl --location 'https://serverless.roboflow.com/infer/lmm' \
  --header 'Content-Type: application/json' \
  --data '{
    "api_key": "'"$ROBOFLOW_API_KEY"'",
    "image": {"type": "url", "value": "https://media.roboflow.com/quickstart/dog.jpeg"},
    "model_id": "florence-2-base",
    "prompt": "<CAPTION>"
  }'

Set api_url to match your deployment target:

  • https://serverless.roboflow.com for the Serverless Cloud API.
  • http://localhost:9001 for a local Inference server.
  • Your Dedicated Deployment URL for a private endpoint.

Swap <CAPTION> for any supported task prompt (for example <DETAILED_CAPTION>, <OD>, <OCR>, <OPEN_VOCABULARY_DETECTION>, <REFERRING_EXPRESSION_SEGMENTATION>) to switch between captioning, detection, OCR, and segmentation tasks.

For self-hosted deployment and the full list of task prompts, see the Inference documentation.

Florence 2 task prompts

Florence 2 switches task by prompt token. Pass one of the following as prompt:

TaskPrompt
Object detection<OD>
Dense region captioning<DENSE_REGION_CAPTION>
Image captioning<CAPTION>, <DETAILED_CAPTION>, <MORE_DETAILED_CAPTION>
Region proposal<REGION_PROPOSAL>
Phrase grounding<CAPTION_TO_PHRASE_GROUNDING>
Referring expression segmentation<REFERRING_EXPRESSION_SEGMENTATION>
Region to segmentation<REGION_TO_SEGMENTATION>
Open vocabulary detection<OPEN_VOCABULARY_DETECTION>
Region to description<REGION_TO_DESCRIPTION>
OCR<OCR>
OCR with region<OCR_WITH_REGION>

Run Florence 2 with self-hosted Inference

Florence 2 can also be loaded directly with the inference package.

1

Install the package

pip install "inference[transformers]"

Use inference-gpu[transformers] on a GPU machine.

2

Run the model

from inference import get_model

model = get_model("florence-2-base", api_key="YOUR_API_KEY")

result = model.infer(
    "https://media.roboflow.com/inference/seawithdock.jpeg",
    prompt="<CAPTION>",
)

print(result[0].response)

Swap <CAPTION> for any task prompt from the table above.

Execution modes in Workflows

When used in a Workflow, Florence 2 runs in one of two modes:

  • Local execution: the model runs on your Inference server (GPU recommended).
  • Remote execution: the model is invoked over HTTP on a remote Inference server.