fal-ai-media
affaan-m/ECC
透過 MCP 工具使用 fal.ai 模型生成圖像、影片和音訊,支援文字轉圖像、文字/圖像轉影片、文字轉語音,以及影片轉音訊等功能。
...展開全部fal.ai 媒體生成
此技能易受變動影響。fal.ai 模型 ID、定價、輸入參數及 MCP 工具名稱 會迅速變更。在承諾使用 特定模型、參數、輸出格式或成本之前,請先搜尋或取得當前模型的元資料。
透過 MCP 使用 fal.ai 模型生成圖片、影片及音訊。
何時啟用
- 使用者希望根據文字提示生成圖片
- 根據文字或圖片生成影片
- 生成語音、音樂或音效
- 任何媒體生成任務
- 使用者說出「生成圖片」、「製作影片」、「文字轉語音」、「製作縮圖」或類似指令
MCP 需求
必須設定 fal.ai MCP 伺服器。新增至 ~/.claude.json:
"fal-ai": {
"command": "npx",
"args": ["-y", "fal-ai-mcp-server"],
"env": { "FAL_KEY": "YOUR_FAL_KEY_HERE" }
}
請至 fal.ai 取得 API 金鑰。
MCP 工具
fal.ai MCP 提供以下工具:
search— 透過關鍵字搜尋可用模型find— 檢視模型詳細資訊與參數generate— 根據參數執行模型result— 檢查非同步生成狀態status— 查看工作狀態cancel— 取消正在執行的任務estimate_cost— 估算生成成本models— 列出熱門模型upload— 上傳檔案作為輸入資料
影像生成
Nano Banana 2 (快速版)
最適合:快速迭代、草稿、文字轉圖像、圖片編輯。
generate(
app_id: "fal-ai/nano-banana-2",
input_data: {
"prompt": "a futuristic cityscape at sunset, cyberpunk style",
"image_size": "landscape_16_9",
"num_images": 1,
"seed": 42
}
)
Nano Banana Pro(高保真)
最適合:正式產出圖像、寫實風格、排版設計、詳細提示詞。
generate(
app_id: "fal-ai/nano-banana-pro",
input_data: {
"prompt": "professional product photo of wireless headphones on marble surface, studio lighting",
"image_size": "square",
"num_images": 1,
"guidance_scale": 7.5
}
)
常見的圖像參數
| 參數 | 類型 | 選項 | 備註 |
|---|---|---|---|
prompt |
字串 | 必填 | 請描述您的需求 |
image_size |
字串 | square, portrait_4_3, landscape_16_9, portrait_16_9, landscape_4_3 |
長寬比 |
num_images |
數字 | 1-4 | 要產生多少個 |
seed |
數字傳統中文(台灣) | 任意整數 | 可重現性 |
guidance_scale |
數字 | 1-20 | 遵循提示的程度(數值越高 = 越字面化) |
影像編輯
使用 Nano Banana 2 搭配輸入圖片進行填補、擴展或風格轉移:
# First upload the source image
upload(file_path: "/path/to/image.png")
# Then generate with image input
generate(
app_id: "fal-ai/nano-banana-2",
input_data: {
"prompt": "same scene but in watercolor style",
"image_url": "",
"image_size": "landscape_16_9"
}
)
影片生成
Seedance 1.0 Pro(字節跳動)
最適合:文字轉影片、圖片轉影片,且具備高畫質動態效果。
generate(
app_id: "fal-ai/seedance-1-0-pro",
input_data: {
"prompt": "a drone flyover of a mountain lake at golden hour, cinematic",
"duration": "5s",
"aspect_ratio": "16:9",
"seed": 42
}
)
Kling Video v3 Pro
最適合:文字/圖片轉影片,並具備原生音訊生成功能。
generate(
app_id: "fal-ai/kling-video/v3/pro",
input_data: {
"prompt": "ocean waves crashing on a rocky coast, dramatic clouds",
"duration": "5s",
"aspect_ratio": "16:9"
}
)
Veo 3(Google DeepMind)
最適合:具備生成音效且視覺品質優異的影片。
generate(
app_id: "fal-ai/veo-3",
input_data: {
"prompt": "a bustling Tokyo street market at night, neon signs, crowd noise",
"aspect_ratio": "16:9"
}
)
圖片轉影片
從現有圖片開始:
generate(
app_id: "fal-ai/seedance-1-0-pro",
input_data: {
"prompt": "camera slowly zooms out, gentle wind moves the trees",
"image_url": "",
"duration": "5s"
}
)
影片參數
| 參數 | 類型 | 選項 | 備註 |
|---|---|---|---|
prompt |
字串 | 必填 | 描述影片 |
duration |
字串 | "5s", "10s" |
影片長度 |
aspect_ratio |
字串 | "16:9", "9:16", "1:1" |
畫面比例 |
seed |
數字 | 任意整數 | 可重現性 |
image_url |
字串 | 網址 | 「圖片轉影片」的原始圖片 |
音訊生成
CSM-1B(對話式語音)
具備自然、對話般品質的文字轉語音功能。
generate(
app_id: "fal-ai/csm-1b",
input_data: {
"text": "Hello, welcome to the demo. Let me show you how this works.",
"speaker_id": 0
}
)
ThinkSound(影片轉音訊)
從影片內容生成相應的音訊。
generate(
app_id: "fal-ai/thinksound",
input_data: {
"video_url": "",
"prompt": "ambient forest sounds with birds chirping"
}
)
ElevenLabs(透過 API,無需 MCP)
如需專業級語音合成服務,請直接使用 ElevenLabs:
import os
import requests
resp = requests.post(
"https://api.elevenlabs.io/v1/text-to-speech/",
headers={
"xi-api-key": os.environ["ELEVENLABS_API_KEY"],
"Content-Type": "application/json"
},
json={
"text": "Your text here",
"model_id": "eleven_turbo_v2_5",
"voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
}
)
with open("output.mp3", "wb") as f:
f.write(resp.content)
VideoDB 生成式音訊
若已設定 VideoDB,請使用其生成式音訊功能:
# Voice generation
audio = coll.generate_voice(text="Your narration here", voice="alloy")
# Music generation
music = coll.generate_music(prompt="upbeat electronic background music", duration=30)
# Sound effects
sfx = coll.generate_sound_effect(prompt="thunder crack followed by rain")
成本估算
生成前,請先確認預估成本:
estimate_cost(
estimate_type: "unit_price",
endpoints: {
"fal-ai/nano-banana-pro": {
"unit_quantity": 1
}
}
)
模型搜尋
尋找適用於特定任務的模型:
search(query: "text to video")
find(endpoint_ids: ["fal-ai/seedance-1-0-pro"])
models()
提示
- 使用
seed可重複的結果來反覆調整提示詞 - 先使用成本較低的模型(Nano Banana 2)進行提示詞迭代,最後再切換至 Pro 版本以完成最終版本
- 針對影片,提示語應具描述性但簡潔——著重於動作與場景
- 相較於純文字轉影片,圖像轉影片能產生更可控的結果
- 請檢查
estimate_cost在執行耗費資源的影片生成前
相關技能
videodb— 影片處理、剪輯與串流video-editing— 由人工智慧驅動的影片編輯工作流程content-engine— 社群平台內容創作
---
name: fal-ai-media
description: Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
---
# fal.ai Media Generation
> **Drift-prone skill.** fal.ai model IDs, pricing, inputs, and MCP tool names
> change quickly. Search or fetch the current model metadata before promising a
> specific model, parameter, output format, or cost.
Generate images, videos, and audio using fal.ai models via MCP.
## When to Activate
- User wants to generate images from text prompts
- Creating videos from text or images
- Generating speech, music, or sound effects
- Any media generation task
- User says "generate image", "create video", "text to speech", "make a thumbnail", or similar
## MCP Requirement
fal.ai MCP server must be configured. Add to `~/.claude.json`:
```json
"fal-ai": {
"command": "npx",
"args": ["-y", "fal-ai-mcp-server"],
"env": { "FAL_KEY": "YOUR_FAL_KEY_HERE" }
}
```
Get an API key at [fal.ai](https://fal.ai).
## MCP Tools
The fal.ai MCP provides these tools:
- `search` — Find available models by keyword
- `find` — Get model details and parameters
- `generate` — Run a model with parameters
- `result` — Check async generation status
- `status` — Check job status
- `cancel` — Cancel a running job
- `estimate_cost` — Estimate generation cost
- `models` — List popular models
- `upload` — Upload files for use as inputs
---
## Image Generation
### Nano Banana 2 (Fast)
Best for: quick iterations, drafts, text-to-image, image editing.
```
generate(
app_id: "fal-ai/nano-banana-2",
input_data: {
"prompt": "a futuristic cityscape at sunset, cyberpunk style",
"image_size": "landscape_16_9",
"num_images": 1,
"seed": 42
}
)
```
### Nano Banana Pro (High Fidelity)
Best for: production images, realism, typography, detailed prompts.
```
generate(
app_id: "fal-ai/nano-banana-pro",
input_data: {
"prompt": "professional product photo of wireless headphones on marble surface, studio lighting",
"image_size": "square",
"num_images": 1,
"guidance_scale": 7.5
}
)
```
### Common Image Parameters
| Param | Type | Options | Notes |
|-------|------|---------|-------|
| `prompt` | string | required | Describe what you want |
| `image_size` | string | `square`, `portrait_4_3`, `landscape_16_9`, `portrait_16_9`, `landscape_4_3` | Aspect ratio |
| `num_images` | number | 1-4 | How many to generate |
| `seed` | number | any integer | Reproducibility |
| `guidance_scale` | number | 1-20 | How closely to follow the prompt (higher = more literal) |
### Image Editing
Use Nano Banana 2 with an input image for inpainting, outpainting, or style transfer:
```
# First upload the source image
upload(file_path: "/path/to/image.png")
# Then generate with image input
generate(
app_id: "fal-ai/nano-banana-2",
input_data: {
"prompt": "same scene but in watercolor style",
"image_url": "<uploaded_url>",
"image_size": "landscape_16_9"
}
)
```
---
## Video Generation
### Seedance 1.0 Pro (ByteDance)
Best for: text-to-video, image-to-video with high motion quality.
```
generate(
app_id: "fal-ai/seedance-1-0-pro",
input_data: {
"prompt": "a drone flyover of a mountain lake at golden hour, cinematic",
"duration": "5s",
"aspect_ratio": "16:9",
"seed": 42
}
)
```
### Kling Video v3 Pro
Best for: text/image-to-video with native audio generation.
```
generate(
app_id: "fal-ai/kling-video/v3/pro",
input_data: {
"prompt": "ocean waves crashing on a rocky coast, dramatic clouds",
"duration": "5s",
"aspect_ratio": "16:9"
}
)
```
### Veo 3 (Google DeepMind)
Best for: video with generated sound, high visual quality.
```
generate(
app_id: "fal-ai/veo-3",
input_data: {
"prompt": "a bustling Tokyo street market at night, neon signs, crowd noise",
"aspect_ratio": "16:9"
}
)
```
### Image-to-Video
Start from an existing image:
```
generate(
app_id: "fal-ai/seedance-1-0-pro",
input_data: {
"prompt": "camera slowly zooms out, gentle wind moves the trees",
"image_url": "<uploaded_image_url>",
"duration": "5s"
}
)
```
### Video Parameters
| Param | Type | Options | Notes |
|-------|------|---------|-------|
| `prompt` | string | required | Describe the video |
| `duration` | string | `"5s"`, `"10s"` | Video length |
| `aspect_ratio` | string | `"16:9"`, `"9:16"`, `"1:1"` | Frame ratio |
| `seed` | number | any integer | Reproducibility |
| `image_url` | string | URL | Source image for image-to-video |
---
## Audio Generation
### CSM-1B (Conversational Speech)
Text-to-speech with natural, conversational quality.
```
generate(
app_id: "fal-ai/csm-1b",
input_data: {
"text": "Hello, welcome to the demo. Let me show you how this works.",
"speaker_id": 0
}
)
```
### ThinkSound (Video-to-Audio)
Generate matching audio from video content.
```
generate(
app_id: "fal-ai/thinksound",
input_data: {
"video_url": "<video_url>",
"prompt": "ambient forest sounds with birds chirping"
}
)
```
### ElevenLabs (via API, no MCP)
For professional voice synthesis, use ElevenLabs directly:
```python
import os
import requests
resp = requests.post(
"https://api.elevenlabs.io/v1/text-to-speech/<voice_id>",
headers={
"xi-api-key": os.environ["ELEVENLABS_API_KEY"],
"Content-Type": "application/json"
},
json={
"text": "Your text here",
"model_id": "eleven_turbo_v2_5",
"voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
}
)
with open("output.mp3", "wb") as f:
f.write(resp.content)
```
### VideoDB Generative Audio
If VideoDB is configured, use its generative audio:
```python
# Voice generation
audio = coll.generate_voice(text="Your narration here", voice="alloy")
# Music generation
music = coll.generate_music(prompt="upbeat electronic background music", duration=30)
# Sound effects
sfx = coll.generate_sound_effect(prompt="thunder crack followed by rain")
```
---
## Cost Estimation
Before generating, check estimated cost:
```
estimate_cost(
estimate_type: "unit_price",
endpoints: {
"fal-ai/nano-banana-pro": {
"unit_quantity": 1
}
}
)
```
## Model Discovery
Find models for specific tasks:
```
search(query: "text to video")
find(endpoint_ids: ["fal-ai/seedance-1-0-pro"])
models()
```
## Tips
- Use `seed` for reproducible results when iterating on prompts
- Start with lower-cost models (Nano Banana 2) for prompt iteration, then switch to Pro for finals
- For video, keep prompts descriptive but concise — focus on motion and scene
- Image-to-video produces more controlled results than pure text-to-video
- Check `estimate_cost` before running expensive video generations
## Related Skills
- `videodb` — Video processing, editing, and streaming
- `video-editing` — AI-powered video editing workflows
- `content-engine` — Content creation for social platforms
所有檔案
1 個檔案安裝 fal-ai-media
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/affaan-m/ECC/tree/main/skills/fal-ai-media # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
