Agnes AI Made Text/Image/Video Completely Free — Here's How I Integrated It into Hermes Agent
An AI platform made text/image/video completely free, and I integrated it into Hermes Agent. Sapiens AI, one of the top 10 AI labs globally, recently did something that made developers call them a godsend: they opened up all three types of multimodal models on their Agnes platform—text, image, and video—for free. Not a limited-time offer, not a trial quota, just a straight-up $0. Someone on Cla
💡 What You Will Learn
An AI platform made text/image/video completely free, and I integrated it into Hermes Agent. Sapiens AI, one of the top 10 AI labs globally, recently did something that made developers call them a god
📜 Table of Contents
- How Complete Is the Agnes Family
- Registration and Getting Your API Key
- Integrating into Hermes Agent Takes Just Three Steps
- Text-to-Image and Image-to-Image: One Sentence for an Image, One Image for a Style Change
- Text-to-Video and Image-to-Video: Text Becomes a Movie, Images Become Short Clips
- Summary
An AI Platform Made Text/Image/Video Completely Free, and I Integrated It into Hermes Agent
Sapiens AI, a top-10 global AI lab, recently did something that made developers call them saints: they opened up all three types of multimodal models on their Agnes platform—text, image, and video—completely free. Not time-limited free, not trial credits, just a straight-up $0.
Some people are pitting it against Claude Opus 4.6 and GLM 5.1 on the Claw-Eval PASS³ benchmark, some are hitting Elo 1178 on the AA image editing leaderboard, and others are using AA's image-to-video Elo 934 for image-to-video conversion—these capabilities now cost zero dollars to use. I spent an afternoon integrating all of them into Hermes Agent, running through everything from text chat to text-to-image, image-to-image, and then text-to-video and image-to-video. Here's the real deal below.
How Complete Is the Agnes Family
Let's start with the model lineup. The text workhorse is Agnes 2.0 Flash, with a 512K context window, supporting tool calling, multi-turn dialogue, and Thinking Mode, scoring 60.9% on Claw-Eval PASS³, beating DeepSeek V4 Pro's 58.4%. On the image side, Agnes Image 2.1 Flash focuses on high information density and composition preservation—give it an image to restyle, and the original layout won't drift. For video, Agnes Video V2.0 supports three modes: text-to-video, image-to-video, and keyframe animation, with async API and async retrieval, up to 18 seconds at 720p.
Standard pricing is $0.03/1M input tokens for text, $0.003/image, and $0.005/second for video—already rock-bottom in the industry. But right now it's $0. You read that right, $0. As for when they'll restore pricing, nobody knows, but jumping on board while it's free is always the right call.
Registration and Getting Your API Key
Go to platform.agnes-ai.com to register an account, verify your email, log in, head to the API Platform page, and click "Create API Key" to generate one in a single click. Copy the key down—you'll need it for configuration later.
Integrating into Hermes Agent Takes Just Three Steps
Hermes Agent supports custom OpenAI-compatible providers, and Agnes's API is fully OpenAI-format compatible, so migration costs zero effort.
Step 1: Add the provider in config.yaml. Windows users, watch the path—the file is at C:\Users\<username>\AppData\Local\hermes\config.yaml. Many people write ~/.hermes instead and end up staring at a 401 error for ages without knowing why.
providers:
agnes:
name: Agnes AI
api: https://apihub.agnes-ai.com/v1
api_key: ${AGNES_CRED}
api_mode: chat_completions
default_model: agnes-2.0-flash
discover_models: true
Step 2: Write AGNES_CRED=your_API_KEY into the .env file in the same directory. Get the API Key by registering at platform.agnes-ai.com.
Step 3: Switch models via the command line and restart the TUI.
hermes config set model.provider custom:agnes
hermes config set model.default agnes-2.0-flash
Done. Hermes is now switched to Agnes 2.0 Flash as the default text model.
Text-to-Image and Image-to-Image: One Sentence for an Image, One Image for a Style Change
Agnes Image 2.1 Flash's API endpoint is POST /v1/images/generations. Text-to-image is a one-liner:
curl -X POST "https://apihub.agnes-ai.com/v1/images/generations" \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"model": "agnes-image-2.1-flash",
"prompt": "A luminous floating city above a misty canyon at sunrise",
"size": "2K",
"ratio": "16:9"
}'
The data[0].url in the response is the image link. I tested it by generating a 2K 16:9 sunrise landscape of Lake Moraine—rich detail, accurate colors, and a 7MB file with solid quality.
Image-to-image is even more interesting. Upload an image, change the style, color, or scene, and the composition stays intact. For example, turning a lakeside landscape into a rain-soaked cyberpunk night:
curl -X POST "https://apihub.agnes-ai.com/v1/images/generations" \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"model": "agnes-image-2.1-flash",
"prompt": "Transform into a rain-soaked cyberpunk night with neon reflections, preserve original composition",
"size": "2K",
"ratio": "16:9",
"extra_body": {
"image": ["https://example.com/input.png"],
"response_format": "url"
}
}'
One gotcha: response_format can't go in the top-level JSON—it must be inside extra_body. Also, the API normalizes your exact pixel values to the nearest preset, so I recommend using the size + ratio combination directly instead of hand-writing width and height.
Text-to-Video and Image-to-Video: Text Becomes a Movie, Images Become Short Clips
The video API is async, with three steps: create a task, poll the status, and download the video.
Create the task:
curl -X POST "https://apihub.agnes-ai.com/v1/videos" \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"model": "agnes-video-v2.0",
"prompt": "A cinematic drone shot over a misty mountain range at sunrise, golden light breaking through clouds, smooth motion, photorealistic",
"width": 1152,
"height": 768,
"num_frames": 121,
"frame_rate": 24
}'
After getting the video_id back, poll with this endpoint:
curl "https://apihub.agnes-ai.com/agnesapi?video_id=<VIDEO_ID>" \
-H "Authorization: Bearer ***"
The status goes from queued to in_progress to completed, and once done, the url field is the video link. I also tested a 5-second aerial mountain video—inference took about 103 seconds, and the final output was reasonably smooth.
Image-to-video works the same way—just add an image parameter with the image URL to the request. For keyframe animation, you'll need to pass multiple image URLs using extra_body.mode: "keyframes".
Video has two easy pitfalls: first, num_frames must be 8n+1 (81, 121, 241, 441)—getting it wrong throws an error immediately; second, inference time isn't short, so setting a polling interval of 10 seconds or more is more reasonable.
Summary
Agnes AI is the most generous free API platform I've seen so far—text, image, and video, full multimodal coverage, all free. Integrating it into Hermes Agent only takes a few lines of config, no extra SDK needed. But there's an even simpler method: throw the official integration docs at hermes and say: "Based on the documentation spec, turn the image generation and video generation API calls into a skill." If you're using Hermes Agent or any OpenAI-compatible AI framework, now's the perfect time to jump on. Why wait until they start charging when you can use it for free now?
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
