옵션

fal-ai-media

affaan-m/ECC affaan-m/ECC

MCP 도구를 통해 fal.ai 모델을 활용하여 이미지, 동영상 및 오디오를 생성할 수 있으며, 텍스트-이미지 변환, 텍스트/이미지-동영상 변환, 텍스트-음성 변환, 동영상-오디오 변환 기능을 지원합니다.

...모든 것을 확장하십시오
0
업데이트 된 시간 2026년 10월 1일

fal.ai 미디어 생성

변동이 잦은 기능입니다. fal.ai 모델 ID, 가격, 입력값 및 MCP 도구 이름은 빠르게 변경될 수 있습니다. 특정 모델, 매개변수, 출력 형식 또는 비용을 확정하기 전에 현재 모델 메타데이터를 검색하거나 확인하십시오.

MCP를 통해 fal.ai 모델을 사용하여 이미지, 동영상 및 오디오를 생성하십시오.

활성화 시점

  • 사용자가 텍스트 프롬프트로 이미지를 생성하고자 할 때
  • 텍스트나 이미지를 기반으로 동영상을 생성할 때
  • 음성, 음악 또는 음향 효과 생성
  • 모든 미디어 생성 작업
  • 사용자가 “이미지 생성”, “동영상 만들기”, “텍스트 음성 변환”, “썸네일 만들기” 또는 이와 유사한 명령을 말할 때

MCP 요구 사항

fal.ai MCP 서버가 구성되어야 합니다. 다음을 추가하십시오. ~/.claude.json:

"fal-ai": {
  "command": "npx",
  "args": ["-y", "fal-ai-mcp-server"],
  "env": { "FAL_KEY": "YOUR_FAL_KEY_HERE" }
}

fal.ai에서 API 키를 발급받으십시오.

MCP 도구

fal.ai MCP는 다음과 같은 도구를 제공합니다:

  • search — 키워드로 사용 가능한 모델 찾기
  • find — 모델 세부 정보 및 매개변수 조회
  • generate — 매개변수를 사용하여 모델 실행
  • result — 비동기 생성 상태 확인
  • status — 작업 상태 확인
  • cancel — 실행 중인 작업 취소
  • estimate_cost — 생성 비용 추정
  • models — 인기 모델 목록 확인
  • upload — 입력 자료로 사용할 파일 업로드

이미지 생성

Nano Banana 2 (고속)

가장 적합한 용도: 빠른 반복 작업, 초안 작성, 텍스트-이미지 변환, 이미지 편집.

generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "a futuristic cityscape at sunset, cyberpunk style",
    "image_size": "landscape_16_9",
    "num_images": 1,
    "seed": 42
  }
)

Nano Banana Pro (고화질)

가장 적합한 용도: 제작용 이미지, 사실적인 표현, 타이포그래피, 상세한 프롬프트.

generate(
  app_id: "fal-ai/nano-banana-pro",
  input_data: {
    "prompt": "professional product photo of wireless headphones on marble surface, studio lighting",
    "image_size": "square",
    "num_images": 1,
    "guidance_scale": 7.5
  }
)

일반적인 이미지 매개변수

매개변수 유형 옵션 비고
prompt 문자열 필수 원하는 내용을 설명해 주세요
image_size 문자열 square, portrait_4_3, landscape_16_9, portrait_16_9, landscape_4_3 화면비
num_images 숫자 1-4 생성할 개수
seed 숫자1-4생성할 개수 임의의 정수 재현성
guidance_scale 숫자 1-20 지시 사항을 얼마나 충실히 따를지 (높을수록 = 더 문자 그대로 따름)

이미지 편집

입력 이미지를 사용하여 Nano Banana 2로 인페인팅, 아웃페인팅 또는 스타일 전이를 수행하세요:

# First upload the source image
upload(file_path: "/path/to/image.png")

# Then generate with image input
generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "same scene but in watercolor style",
    "image_url": "",
    "image_size": "landscape_16_9"
  }
)

동영상 생성

Seedance 1.0 Pro (ByteDance)

가장 적합한 용도: 텍스트-비디오 변환, 움직임 품질이 뛰어난 이미지-비디오 변환.

generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "a drone flyover of a mountain lake at golden hour, cinematic",
    "duration": "5s",
    "aspect_ratio": "16:9",
    "seed": 42
  }
)

Kling Video v3 Pro

가장 적합한 용도: 네이티브 오디오 생성이 가능한 텍스트/이미지-비디오 변환.

generate(
  app_id: "fal-ai/kling-video/v3/pro",
  input_data: {
    "prompt": "ocean waves crashing on a rocky coast, dramatic clouds",
    "duration": "5s",
    "aspect_ratio": "16:9"
  }
)

Veo 3 (Google DeepMind)

가장 적합한 용도: 생성된 사운드와 높은 화질을 갖춘 동영상 제작.

generate(
  app_id: "fal-ai/veo-3",
  input_data: {
    "prompt": "a bustling Tokyo street market at night, neon signs, crowd noise",
    "aspect_ratio": "16:9"
  }
)

이미지-동영상 변환

기존 이미지를 기반으로 시작:

generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "camera slowly zooms out, gentle wind moves the trees",
    "image_url": "",
    "duration": "5s"
  }
)

동영상 매개변수

매개변수 유형 옵션 참고
prompt 문자열 필수 동영상 설명
duration 문자열 "5s", "10s" 동영상 길이
aspect_ratio 문자열 "16:9", "9:16", "1:1" 프레임 비율
seed 숫자 임의의 정수 재현성
image_url 문자열 URL 이미지-동영상 변환용 원본 이미지

오디오 생성

CSM-1B (대화형 음성)

자연스럽고 대화체 같은 음성의 텍스트-음성 변환.

generate(
  app_id: "fal-ai/csm-1b",
  input_data: {
    "text": "Hello, welcome to the demo. Let me show you how this works.",
    "speaker_id": 0
  }
)

ThinkSound (영상-오디오 변환)

동영상 콘텐츠에서 일치하는 오디오를 생성합니다.

generate(
  app_id: "fal-ai/thinksound",
  input_data: {
    "video_url": "",
    "prompt": "ambient forest sounds with birds chirping"
  }
)

ElevenLabs (API를 통해, MCP 불필요)

전문적인 음성 합성을 위해서는 ElevenLabs를 직접 사용하세요:

import os
import requests

resp = requests.post(
    "https://api.elevenlabs.io/v1/text-to-speech/",
    headers={
        "xi-api-key": os.environ["ELEVENLABS_API_KEY"],
        "Content-Type": "application/json"
    },
    json={
        "text": "Your text here",
        "model_id": "eleven_turbo_v2_5",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }
)
with open("output.mp3", "wb") as f:
    f.write(resp.content)

VideoDB 생성형 오디오

VideoDB가 설정되어 있다면, 해당 생성형 오디오 기능을 사용하세요:

# Voice generation
audio = coll.generate_voice(text="Your narration here", voice="alloy")

# Music generation
music = coll.generate_music(prompt="upbeat electronic background music", duration=30)

# Sound effects
sfx = coll.generate_sound_effect(prompt="thunder crack followed by rain")

비용 추정

생성하기 전에 예상 비용을 확인하세요:

estimate_cost(
  estimate_type: "unit_price",
  endpoints: {
    "fal-ai/nano-banana-pro": {
      "unit_quantity": 1
    }
  }
)

모델 검색

특정 작업에 적합한 모델을 찾으세요:

search(query: "text to video")
find(endpoint_ids: ["fal-ai/seedance-1-0-pro"])
models()

팁

  • 다음과 같이 사용하십시오 seed 프롬프트 반복 작업 시 재현 가능한 결과를 얻으려면
  • 프롬프트 반복 작업 시에는 비용이 저렴한 모델(Nano Banana 2)로 시작하고, 최종 결과물 작업 시에는 Pro 모델로 전환하세요
  • 동영상의 경우, 프롬프트는 설명적이면서도 간결하게 작성하세요 — 동작과 장면에 중점을 두세요
  • 이미지-투-비디오 방식은 순수한 텍스트-투-비디오 방식보다 결과물을 더 정교하게 제어할 수 있습니다
  • 확인 estimate_cost 비용이 많이 드는 동영상 생성을 실행하기 전에

관련 기술

  • videodb — 동영상 처리, 편집 및 스트리밍
  • video-editing — AI 기반 동영상 편집 워크플로우
  • content-engine — 소셜 플랫폼용 콘텐츠 제작
GitHub에서 보기
---
name: fal-ai-media
description: Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
---

# fal.ai Media Generation

> **Drift-prone skill.** fal.ai model IDs, pricing, inputs, and MCP tool names
> change quickly. Search or fetch the current model metadata before promising a
> specific model, parameter, output format, or cost.

Generate images, videos, and audio using fal.ai models via MCP.

## When to Activate

- User wants to generate images from text prompts
- Creating videos from text or images
- Generating speech, music, or sound effects
- Any media generation task
- User says "generate image", "create video", "text to speech", "make a thumbnail", or similar

## MCP Requirement

fal.ai MCP server must be configured. Add to `~/.claude.json`:

```json
"fal-ai": {
  "command": "npx",
  "args": ["-y", "fal-ai-mcp-server"],
  "env": { "FAL_KEY": "YOUR_FAL_KEY_HERE" }
}
```

Get an API key at [fal.ai](https://fal.ai).

## MCP Tools

The fal.ai MCP provides these tools:
- `search` — Find available models by keyword
- `find` — Get model details and parameters
- `generate` — Run a model with parameters
- `result` — Check async generation status
- `status` — Check job status
- `cancel` — Cancel a running job
- `estimate_cost` — Estimate generation cost
- `models` — List popular models
- `upload` — Upload files for use as inputs

---

## Image Generation

### Nano Banana 2 (Fast)
Best for: quick iterations, drafts, text-to-image, image editing.

```
generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "a futuristic cityscape at sunset, cyberpunk style",
    "image_size": "landscape_16_9",
    "num_images": 1,
    "seed": 42
  }
)
```

### Nano Banana Pro (High Fidelity)
Best for: production images, realism, typography, detailed prompts.

```
generate(
  app_id: "fal-ai/nano-banana-pro",
  input_data: {
    "prompt": "professional product photo of wireless headphones on marble surface, studio lighting",
    "image_size": "square",
    "num_images": 1,
    "guidance_scale": 7.5
  }
)
```

### Common Image Parameters

| Param | Type | Options | Notes |
|-------|------|---------|-------|
| `prompt` | string | required | Describe what you want |
| `image_size` | string | `square`, `portrait_4_3`, `landscape_16_9`, `portrait_16_9`, `landscape_4_3` | Aspect ratio |
| `num_images` | number | 1-4 | How many to generate |
| `seed` | number | any integer | Reproducibility |
| `guidance_scale` | number | 1-20 | How closely to follow the prompt (higher = more literal) |

### Image Editing
Use Nano Banana 2 with an input image for inpainting, outpainting, or style transfer:

```
# First upload the source image
upload(file_path: "/path/to/image.png")

# Then generate with image input
generate(
  app_id: "fal-ai/nano-banana-2",
  input_data: {
    "prompt": "same scene but in watercolor style",
    "image_url": "<uploaded_url>",
    "image_size": "landscape_16_9"
  }
)
```

---

## Video Generation

### Seedance 1.0 Pro (ByteDance)
Best for: text-to-video, image-to-video with high motion quality.

```
generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "a drone flyover of a mountain lake at golden hour, cinematic",
    "duration": "5s",
    "aspect_ratio": "16:9",
    "seed": 42
  }
)
```

### Kling Video v3 Pro
Best for: text/image-to-video with native audio generation.

```
generate(
  app_id: "fal-ai/kling-video/v3/pro",
  input_data: {
    "prompt": "ocean waves crashing on a rocky coast, dramatic clouds",
    "duration": "5s",
    "aspect_ratio": "16:9"
  }
)
```

### Veo 3 (Google DeepMind)
Best for: video with generated sound, high visual quality.

```
generate(
  app_id: "fal-ai/veo-3",
  input_data: {
    "prompt": "a bustling Tokyo street market at night, neon signs, crowd noise",
    "aspect_ratio": "16:9"
  }
)
```

### Image-to-Video
Start from an existing image:

```
generate(
  app_id: "fal-ai/seedance-1-0-pro",
  input_data: {
    "prompt": "camera slowly zooms out, gentle wind moves the trees",
    "image_url": "<uploaded_image_url>",
    "duration": "5s"
  }
)
```

### Video Parameters

| Param | Type | Options | Notes |
|-------|------|---------|-------|
| `prompt` | string | required | Describe the video |
| `duration` | string | `"5s"`, `"10s"` | Video length |
| `aspect_ratio` | string | `"16:9"`, `"9:16"`, `"1:1"` | Frame ratio |
| `seed` | number | any integer | Reproducibility |
| `image_url` | string | URL | Source image for image-to-video |

---

## Audio Generation

### CSM-1B (Conversational Speech)
Text-to-speech with natural, conversational quality.

```
generate(
  app_id: "fal-ai/csm-1b",
  input_data: {
    "text": "Hello, welcome to the demo. Let me show you how this works.",
    "speaker_id": 0
  }
)
```

### ThinkSound (Video-to-Audio)
Generate matching audio from video content.

```
generate(
  app_id: "fal-ai/thinksound",
  input_data: {
    "video_url": "<video_url>",
    "prompt": "ambient forest sounds with birds chirping"
  }
)
```

### ElevenLabs (via API, no MCP)
For professional voice synthesis, use ElevenLabs directly:

```python
import os
import requests

resp = requests.post(
    "https://api.elevenlabs.io/v1/text-to-speech/<voice_id>",
    headers={
        "xi-api-key": os.environ["ELEVENLABS_API_KEY"],
        "Content-Type": "application/json"
    },
    json={
        "text": "Your text here",
        "model_id": "eleven_turbo_v2_5",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }
)
with open("output.mp3", "wb") as f:
    f.write(resp.content)
```

### VideoDB Generative Audio
If VideoDB is configured, use its generative audio:

```python
# Voice generation
audio = coll.generate_voice(text="Your narration here", voice="alloy")

# Music generation
music = coll.generate_music(prompt="upbeat electronic background music", duration=30)

# Sound effects
sfx = coll.generate_sound_effect(prompt="thunder crack followed by rain")
```

---

## Cost Estimation

Before generating, check estimated cost:

```
estimate_cost(
  estimate_type: "unit_price",
  endpoints: {
    "fal-ai/nano-banana-pro": {
      "unit_quantity": 1
    }
  }
)
```

## Model Discovery

Find models for specific tasks:

```
search(query: "text to video")
find(endpoint_ids: ["fal-ai/seedance-1-0-pro"])
models()
```

## Tips

- Use `seed` for reproducible results when iterating on prompts
- Start with lower-cost models (Nano Banana 2) for prompt iteration, then switch to Pro for finals
- For video, keep prompts descriptive but concise — focus on motion and scene
- Image-to-video produces more controlled results than pure text-to-video
- Check `estimate_cost` before running expensive video generations

## Related Skills

- `videodb` — Video processing, editing, and streaming
- `video-editing` — AI-powered video editing workflows
- `content-engine` — Content creation for social platforms

모든 파일

1개 파일

fal-ai-media 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/affaan-m/ECC/tree/main/skills/fal-ai-media # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용합니다.
저장소 affaan-m/ECC

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR