Instructions to use noamrot/FuseCap_Image_Captioning with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use noamrot/FuseCap_Image_Captioning with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("image-to-text", model="noamrot/FuseCap_Image_Captioning")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("noamrot/FuseCap_Image_Captioning") model = AutoModelForMultimodalLM.from_pretrained("noamrot/FuseCap_Image_Captioning", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from noamrot/FuseCap_Image_Captioning: direct link, hf CLI and curl.
- Browser
- Download file 2.15 kB
-
https://huggingface.co/noamrot/FuseCap_Image_Captioning/resolve/main/README.md
- Command line
-
hf download hf://noamrot/FuseCap_Image_Captioning/README.md
-
curl -L -o README.md https://huggingface.co/noamrot/FuseCap_Image_Captioning/resolve/main/README.md
2.15 kB
metadata
license: mit
inference: false
pipeline_tag: image-to-text
tags:
- image-captioning
FuseCap: Leveraging Large Language Models for Enriched Fused Image Captions
A framework designed to generate semantically rich image captions.
Resources
๐ป Project Page: For more details, visit the official project page.
๐ Read the Paper: You can find the paper here.
๐ Demo: Try out our BLIP-based model demo trained using FuseCap.
๐ Code Repository: The code for FuseCap can be found in the GitHub repository.
๐๏ธ Datasets: The fused captions datasets can be accessed from here.
Running the model
Our BLIP-based model can be run using the following code,
import requests
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration
import torch
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
processor = BlipProcessor.from_pretrained("noamrot/FuseCap")
model = BlipForConditionalGeneration.from_pretrained("noamrot/FuseCap").to(device)
img_url = 'https://huggingface.co/spaces/noamrot/FuseCap/resolve/main/bike.jpg'
raw_image = Image.open(requests.get(img_url, stream=True).raw).convert('RGB')
text = "a picture of "
inputs = processor(raw_image, text, return_tensors="pt").to(device)
out = model.generate(**inputs, num_beams = 3)
print(processor.decode(out[0], skip_special_tokens=True))
Upcoming Updates
The official codebase, datasets and trained models for this project will be released soon.
BibTeX
@inproceedings{rotstein2024fusecap,
title={Fusecap: Leveraging large language models for enriched fused image captions},
author={Rotstein, Noam and Bensa{\"\i}d, David and Brody, Shaked and Ganz, Roy and Kimmel, Ron},
booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision},
pages={5689--5700},
year={2024}
}