Sphinx AI Cloud
API công khai

ASR & TTS API

Bộ API độc lập với SaaS, xác thực bằng API key riêng — thiết kế theo đúng cú pháp ElevenLabs (xi-api-key, /v1/speech-to-text, /v1/text-to-speech) để đổi từ ElevenLabs sang chỉ cần đổi base_url.

API key hiện do admin cấp thủ công (chưa có tự tạo trong dashboard). Key chỉ lưu trong trình duyệt của bạn (localStorage), không gửi đi đâu ngoài các request thử bên dưới.

Base URL: https://market-applications.svisor.vn/public-api · WebSocket: wss://market-applications.svisor.vn/public-api

Mọi request lỗi trả về

{
  "detail": {
    "status": "invalid_api_key",
    "message": "the provided API key is invalid or has been revoked"
  }
}

Voices

GET/v1/voices

Danh sách voice hệ thống và voice bạn đã tự clone.

cURL
curl https://market-applications.svisor.vn/public-api/v1/voices \
  -H "xi-api-key: YOUR_API_KEY"
POST/v1/voices/add

Clone 1 voice mới từ file âm thanh mẫu.

cURL
curl -X POST https://market-applications.svisor.vn/public-api/v1/voices/add \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "name=Giọng của tôi" \
  -F "files=@clip.wav"

Text to Speech

POST/v1/text-to-speech/{voice_id}

Tạo giọng nói từ văn bản, trả về trọn file audio.

cURL
curl -X POST https://market-applications.svisor.vn/public-api/v1/text-to-speech/{voice_id} \
  -H "xi-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Xin chào, đây là bài kiểm tra giọng nói."}' \
  --output speech.mp3
POST/v1/text-to-speech/{voice_id}/stream

Giống trên nhưng stream theo từng câu — xem thời gian byte đầu tiên.

cURL
curl -X POST https://market-applications.svisor.vn/public-api/v1/text-to-speech/{voice_id}/stream \
  -H "xi-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Xin chào, đây là bài kiểm tra giọng nói."}' \
  --output speech.mp3

Speech to Text

POST/v1/speech-to-text

Chuyển file âm thanh (WAV khuyến nghị) thành văn bản, xử lý đồng bộ.

cURL
curl -X POST https://market-applications.svisor.vn/public-api/v1/speech-to-text \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "file=@clip.wav" \
  -F "language_code=vi"
GET/v1/speech-to-text/stream

WebSocket, live caption từ mic — cắt câu tự động khi lặng 300ms.

Chưa kết nối
JavaScript
const ws = new WebSocket(
  `wss://market-applications.svisor.vn/public-api/v1/speech-to-text/stream?xi_api_key=${API_KEY}`
);
ws.binaryType = "arraybuffer";
ws.send(JSON.stringify({ type: "session.update", session: { language: "vi" } }));
ws.onmessage = (e) => console.log(JSON.parse(e.data));
// gửi PCM16LE mono 16kHz, mỗi frame 1 lần ws.send(chunk)
// mỗi transcription.segment (is_final=false) thay thế segment cùng segment_id trước đó