Caption Generator from Photo Context
Free to download on every platform. Comes pre-installed on BotFone, BotPad and BotFlip — with extra free apps included.
About this app
WHAT IT DOES
Caption Generator from Photo Context reads the current page context — including image alt text, page title, surrounding text, and any visible image description — and sends it to the AI to generate 3 to 5 social media captions. You can choose from five styles: funny, inspirational, professional, casual, or storytelling. The tool also offers options to include hashtags and emojis in the generated captions. The captions are presented in the panel for review, and you can copy individual captions or all of them at once. A manual input field is also provided for pasting image descriptions or context directly if the automatic extraction doesn't capture what you need.
WHERE IT RUNS
This userscript runs inside major social media and professional platforms including Instagram, Facebook, X (Twitter), Reddit, and LinkedIn in your browser, powered by the free BotGentz extension. It operates entirely on the page you are viewing — it reads the page context and sends it to the AI backend for caption generation. All AI processing happens remotely, but your page context is only sent for the purpose of generating captions.
HOW TO USE
Install the BotGentz extension and ensure you have an active subscription, then install this userscript from the BotGentz catalogue. Navigate to a page containing a photo, image, or post — the tool will automatically extract context such as alt text, page title, and surrounding text. Select your desired caption style from the dropdown (funny, inspirational, professional, casual, or storytelling), choose the number of captions (3 or 5), and toggle hashtag and emoji inclusion. Click "Generate Captions" to produce your captions. Review the captions in the panel, then click "Copy All" to copy all captions, "Copy First Caption" to copy just the first one, or "Regenerate" to try a different set. You can also paste custom context into the manual input field if the automatic extraction misses important details.
HOW IT WORKS UNDER THE HOOD
The script first attempts to extract page context using WebMCP if available — the new W3C proposal for structured browser tools — which allows direct, typed access to page context without DOM scraping. If WebMCP is not available, the script falls back to a multi-strategy DOM extraction approach. It collects image alt text from img elements, page title, meta description, Open Graph data (og:title and og:description), heading text, and visible post or body content using platform-specific selectors. All collected context is combined into a single text input and sent to the AI backend via a secure API call with your subscription key. The AI processes the input with a system prompt that instructs it to generate the requested number of captions in the specified style, optionally including hashtags and emojis, and caps the output at 1,024 tokens. The response is presented in the panel for review.
THE PANEL
The tool runs in the BotGentz panel, a draggable floating interface that stays on top of the page. You can drag the panel by its header, collapse it to a small bubble, and resize it. The panel remembers its position per site, so it stays where you left it when you return to the platform. Press Escape to close the panel when you are done. The panel includes settings for style, count, hashtags, and emojis, a display area showing the generated captions, and action buttons for generating, copying, regenerating, and saving settings.
PLEASE NOTE
This tool requires the free BotGentz extension and an active subscription to the AI service. AI-generated captions should be reviewed before use — they are drafts, not final responses. The tool works only on the current page and may not capture all context depending on platform structure. Social media platforms frequently update their interfaces — if the tool stops working, check for an updated version or report the issue. Your page context is only sent to the AI backend for the purpose of generating captions; no data is stored permanently. The manual input field can be used as a fallback for unsupported platforms or when automatic extraction is incomplete.