업데이트: 2026-06-06
텍스트 음성 변환(TTS)
from openai import OpenAI
client = OpenAI(api_key="sk-xxx", base_url="https://api.crazyrouter.com/v1")
# 음성 생성
response = client.audio.speech.create(
model="tts-1",
voice="alloy",
input="안녕하세요, Crazyrouter API 서비스를 이용해 주셔서 감사합니다."
)
# MP3 파일로 저장
response.stream_to_file("output.mp3")
print("음성이 output.mp3에 저장되었습니다")
사용 가능한 음성
| 음성 | 특징 |
|---|---|
alloy | 중성적, 균형 잡힘 |
echo | 남성, 차분함 |
fable | 남성, 따뜻함 |
onyx | 남성, 중후함 |
nova | 여성, 활기참 |
shimmer | 여성, 부드러움 |
고품질 TTS
# tts-1-hd를 사용하여 더 높은 음질 확보
response = client.audio.speech.create(
model="tts-1-hd-1106",
voice="nova",
input="고품질 음성 합성 예시",
response_format="opus", # mp3, opus, aac, flac 지원
speed=1.0 # 0.25에서 4.0까지
)
response.stream_to_file("output_hd.opus")
음성 텍스트 변환(STT)
# Whisper 음성 인식
audio_file = open("recording.mp3", "rb")
transcript = client.audio.transcriptions.create(
model="whisper-1",
file=audio_file,
language="zh" # 선택 사항, 언어를 지정하면 정확도가 향상됩니다
)
print(transcript.text)
타임스탬프가 포함된 전사
transcript = client.audio.transcriptions.create(
model="whisper-1",
file=open("recording.mp3", "rb"),
response_format="verbose_json",
timestamp_granularities=["segment"]
)
for segment in transcript.segments:
print(f"[{segment['start']:.1f}s - {segment['end']:.1f}s] {segment['text']}")
음성 번역
영어가 아닌 음성을 영어 텍스트로 번역합니다:audio_file = open("chinese_audio.mp3", "rb")
translation = client.audio.translations.create(
model="whisper-1",
file=audio_file
)
print(translation.text) # 영어 출력
TTS가 지원하는 모델은
tts-1, tts-1-hd-1106, tts-1을 포함합니다. STT는 whisper-1 모델을 사용합니다.