What is Fish Audio?
Hanabi AI Inc., a Delaware corporation named in the Fish Audio terms of use, operates Fish Audio as a browser workspace plus a developer API for synthetic speech.

Fish Audio turns scripts into expressive speech with inline emotion tags such as [whispering] and [sobbing], clones voices from short samples, and transcribes audio with speaker and emotion labels. Creators, audiobook narrators, game studios, and API developers use Fish Audio for its S2.1 Pro model and library of 2,000,000+ community
Verified facts from the vendor site · Last verified Aug 21, 2026. Independent Fish Audio review covering pricing, features, who it is for, and alternatives.
Hanabi AI Inc., a Delaware corporation named in the Fish Audio terms of use, operates Fish Audio as a browser workspace plus a developer API for synthetic speech.
Fish Audio is freemium. It starts at $0/month. Paid plans include Free Tier: monthly $0/mo: yearly $0/yr ($0/mo effective); Plus: monthly $15/mo: yearly $66/yr ($5.5/mo effective, save 63%); Pro: monthly $100/mo: yearly $450/yr ($37.5/mo effective, save 63%); Max: monthly $999/mo: yearly $8988/yr ($749/mo effective, save 25%). Check fish.audio for current prices.
Fish Audio has a free tier. Paid plans start at $15/month.
Fish Audio is best for video narrators, audiobook producers, and game studios who need line-by-line emotion control in synthetic speech instead of one flat voice setting across a whole script.
Fish Audio is best for video narrators, audiobook producers, and game studios who need line-by-line emotion control in synthetic speech instead of one flat voice setting across a whole script. Skip it if vendor contradiction: the Free Tier plan card lists "Commercial use" while the homepage FAQ says the free plan is personal use only and the terms restrict unpaid use to personal, non-commercial use.
The main Fish Audio features are S2.1 Pro real-time text-to-speech model, Inline emotion tags: [angry], [sad], [embarrassed], [emphasis], [whispering], [soft], [breathy], [excited], Special performance tags: [laughing], [chuckling], [clear throat], [sobbing], [crying loudly], [sighing], [panting], [groaning], [crowd laughing], [pause], [long pause], Voice cloning advertised at 15 seconds on the API section and as little as 10 seconds in the FAQ, and Speech-to-text with multispeaker labels, emotion tags, and natural language description.
The closest Fish Audio alternatives on Citeware are Wave.video, VEED, Vidnoz AI, HeyGen, PixVerse. Full list: https://citeware.io/alternatives/fish-audio.
Limitations called out on this listing: Vendor contradiction: the Free Tier plan card lists "Commercial use" while the homepage FAQ says the free plan is personal use only and the terms restrict unpaid use to personal, non-commercial use; Vendor contradiction: the annual toggle promotes "3 months off + anniversary 50% off when billed annually" but the billing FAQ says yearly billing saves 33% versus monthly; Vendor contradiction: cloning is advertised at 15 seconds in the API section and as little as 10 seconds in the FAQ; Vendor contradiction on languages: 30+ languages in the API section, 8 named languages in the FAQ, and 8 languages in the meta description; Unused minutes and credits do not roll over; monthly quotas reset at the start of each billing cycle; Free Tier caps each generation at 500 characters, too short for long-form narration; Credit math is indirect, with roughly 600 to 625 credits consumed per minute of generation; Professional voice slots are rationed by tier (1 on Plus, 5 on Pro, 15 on Max).
Fish Audio is available on Web browser, REST API, On-premise deployment (Enterprise).
Fish Audio integrates with Fish Audio REST API, Fish Audio SDKs for developers, Stripe (payment processing), Google sign-in via third-party account credentials, Voice Agent API for conversational applications, GitHub (fishaudio open-source repositories).
vs
Fish Audio vs Wave.videoOnline video editor, live multistreaming studio, recorder, and hosting in one browser platform
vs
Fish Audio vs VEEDGenerate, edit, subtitle and dub videos in one browser workflow
vs
Fish Audio vs Vidnoz AIFree AI video generator with 1900+ avatars, 2000+ voices, and 2800+ templates
vs
Fish Audio vs HeyGenAI avatar video creation, voice cloning, and lip-synced translation in 175+ languages
vs
Fish Audio vs PixVerseProprietary AI video models for clip generation, cinematic production, and real-time interactive worlds
vs
Fish Audio vs MoonvalleyStudio-grade AI video generation powered by Marey, trained on licensed data.Public threads about Fish Audio
Community observations
No recent Hacker News threads
Nothing matched Fish Audio on Hacker News right now. Search X, Reddit, or another network above.
Citeware analysis
Hanabi AI Inc., a Delaware corporation named in the Fish Audio terms of use, operates Fish Audio as a browser workspace plus a developer API for synthetic speech. The Fish Audio homepage headline calls S2.1 Pro "the most expressive, emotionally controllable real-time voice model" and frames the pitch as endurance rather than a demo trick: "Five seconds is easy. Five minutes is the test."
The most distinctive part of Fish Audio is its tag vocabulary. A writer working in the Fish Audio editor can insert emotion tags including [angry], [sad], [embarrassed], [emphasis], [whispering], [soft], [breathy], and [excited], plus special performance tags such as [laughing], [chuckling], [clear throat], [sobbing], [crying loudly], [sighing], [panting], [groaning], [crowd laughing], [background laughter], [pause], and [long pause]. Because those tags sit inline, a Fish Audio script can shift delivery mid-paragraph instead of only between separate renders.
The Fish Audio interface is organized around three tabs shown on the landing page: Text To Speech, Voice Cloning, and Speech To Text. The public demo accepts input against a 30,000-character counter, lets a visitor pick a voice such as Sarah or Adrian from the 2,000,000+ voice library, and renders audio before signup. Fish Audio also markets an end-to-end Voice Agent solution, transcription that includes multispeaker and emotion tags with natural language description, and multilingual output.
Published Fish Audio numbers do not always agree. The API section of the Fish Audio homepage advertises cloning "with perfect fidelity in 15 seconds" while the Fish Audio FAQ says cloning needs as little as 10 seconds of audio. Language coverage is similarly split: the Fish Audio API section says "Speak 30+ languages with any voice" while the Fish Audio FAQ lists English, Japanese, Korean, Chinese, French, German, Arabic, and Spanish, and the meta description says 8 languages.
Fish Audio meters usage in monthly credits, and the Fish Audio pricing FAQ states each minute of generation costs roughly 600 to 625 credits. The Fish Audio Free Tier is $0 per month with no credit card, 8,000 credits monthly, a 500-character cap per generation, 3 public voice slots, standard generation speed, and enhanced voice cloning. Plus on Fish Audio is $15 month-to-month or $66 for a year (displayed as $5.5/mo) with 250,000 credits, up to 200 minutes, 15,000 characters per generation, 10 private voice slots, Voice Design access, and 1 professional voice slot.
Larger Fish Audio workloads step up to Pro at $100 monthly or $450 billed annually (displayed as $37.5/mo) with 2,000,000 credits, up to 1,620 minutes, 3 team seats, 30,000 characters per generation, unlimited voice slots, 5 professional voice slots, and a 7-day money-back guarantee. Fish Audio Max is $999 monthly or $8,988 billed annually (displayed as $749/mo) with 25,000,000 credits, up to 6,250 minutes, 10 team seats, and 15 professional voice slots. Fish Audio Enterprise is custom volume pricing billed annually with zero data retention, on-premise deployment, SOC2 compliance, organization-level pay-as-you-go controls, and SSO marked coming soon.
Procurement teams should note the contradictions in Fish Audio's own fine print. The Fish Audio Free Tier plan card lists "Commercial use" as a feature, yet the Fish Audio homepage FAQ says the free plan is personal use only and that monetizing on YouTube, podcasts, or in business requires a paid upgrade, and the Fish Audio terms restrict unpaid use to internal, personal, non-commercial use. The Fish Audio annual toggle also promotes "3 months off + anniversary 50% off when billed annually" while the Fish Audio billing FAQ says yearly billing saves 33% compared with monthly. Fish Audio further confirms that monthly quotas reset each billing cycle and unused minutes do not roll over.
Against peers in expressive speech synthesis, Fish Audio leans on three levers: an open-source lineage visible on the fishaudio GitHub organization, a very large user-contributed voice library, and per-sentence tag control rather than a single global style slider. Fish Audio marketing copy directly names ElevenLabs in its FAQ and claims "more affordable pricing with comparable quality," which is a vendor assertion rather than a measured result, so buyers evaluating Fish Audio should test the same script on both engines before committing to an annual Fish Audio plan.
Suitable for
Video narrators, audiobook producers, and game studios who need line-by-line emotion control in synthetic speech instead of one flat voice setting across a whole script
Evidence: Homepage (Primary) · Pricing (Primary) · Verified Aug 21, 2026.
Citeware analysis
Evidence: Homepage (Primary) · Verified Aug 21, 2026.
Verified facts
Evidence: Homepage (Primary) · Verified Aug 21, 2026.
Verified facts
5 plans. Scroll to compare.
Evidence: Homepage (Primary) · Pricing (Primary) · Verified Aug 21, 2026.
Verified facts
The company behind Fish Audio is Hanabi AI Inc.
Provides the Fish.Audio platform and related products and services (collectively, the "Services"), operated by Hanabi AI Inc. Users may create a Fish.Audio Account with a username and password, or sign in via a third-party account such as Google. Not responsible for a user's use of the Services in a way that breaks the law. Fish Audio also lists other losses and interruptions it does not cover.
Fish Audio says it collects account and identity, contact details, payment information, product usage and analytics, and device and ip data. That data is used for creating and managing your account or other user profiles.
Address: 1111B Governors Ave STE 48109, Dover, DE 19904, United States
Provides the Fish.Audio platform and related products and services (collectively, the "Services"), operated by Hanabi AI Inc. Users may create a Fish.Audio Account with a username and password, or sign in via a third-party account such as Google. Not responsible for a user's use of the Services in a way that breaks the law. Fish Audio also lists other losses and interruptions it does not cover.
Fish Audio says it collects account and identity, contact details, payment information, product usage and analytics, and device and ip data. That data is used for creating and managing your account or other user profiles.
Account and identity
CollectingCreating and managing your account or other user profiles
Contact details
CollectingCorresponding with you; sending emails and other communications
Payment information
CollectingProcessing orders or other transactions; billing (shared with payment processor Stripe)
Product usage and analytics
CollectingWeb analytics; improving the Services, including testing, research, internal analytics
Device and IP data
CollectingDevice/IP data shared with advertising and analytics partners; fraud protection, security and debugging
User content and uploads
CollectingUser Content from input, file uploads or feedback used to provide the Services
Cookies and tracking
CollectingTracking Tools, Advertising and Opt-Out; collected automatically through Cookies
Location
CollectingGeolocation data (IP-based location, browser time zone) shared with service, advertising and analytics partners
Shared with third parties
CollectingReceived from vendors, analytics providers, advertising partners and social networks; disclosed to service providers and business partners
Advertising and remarketing
CollectingMarketing and selling the Services; showing you advertisements, including interest-based or online behavioral advertising
Used to train AI models
UnclearNot addressed; policy mentions improving the Services including research and development
How long data is kept
UnclearNot specified in the provided text
Evidence: Terms (Primary) · Privacy (Primary) · Verified Aug 21, 2026.
Citeware analysis
Named models, written separately. These are labeled Citeware analysis, not vendor claims.
anthropic · claude-opus-5 · Aug 21, 2026
Fish Audio's inline emotion and performance tags give directors something most TTS engines still withhold: line-by-line control over delivery rather than one global voice setting. The engineering looks stronger than the marketing copy, which contradicts itself on commercial rights, clone length, and language count.
As of August 2026, the tag-level emotion control is the real draw. Verify the free-plan commercial terms in writing before you ship anything with it.
The tag system is the product. Being able to drop [whispering], [sighing], or [long pause] mid-script means you direct a performance instead of re-rendering a whole take to nudge one line.
For audiobook and character work that difference compounds fast, and the public homepage demo lets you hear it before signing up.
Strengths
Watch-outs
openai · gpt-5.6-terra · Aug 21, 2026
Fish Audio is a compelling expressive TTS choice whose inline performance control is more useful than a huge voice catalog alone. It is easy to recommend for experimentation and directed narration, but production buyers should get commercial rights and language coverage clarified in writing.
As of August 2026, Fish Audio stands out for script-level delivery control, though conflicting plan and capability claims weaken buying confidence.
Fish Audio’s best differentiator is control inside the script. Emotion, pause, and performance tags give a narrator or character producer a practical way to shape individual lines without splitting every take into separate jobs.
The free demo also makes auditioning its sound unusually low-friction.
Strengths
Watch-outs
Verified facts
Common questions about Fish Audio.
Hanabi AI Inc., a Delaware corporation named in the Fish Audio terms of use, operates Fish Audio as a browser workspace plus a developer API for synthetic speech. Source: https://fish.audio/
Fish Audio is freemium. It starts at $0/month. Paid plans include Free Tier: monthly $0/mo: yearly $0/yr ($0/mo effective); Plus: monthly $15/mo: yearly $66/yr ($5.5/mo effective, save 63%); Pro: monthly $100/mo: yearly $450/yr ($37.5/mo effective, save 63%); Max: monthly $999/mo: yearly $8988/yr ($749/mo effective, save 25%). Check fish.audio for current prices. Source: https://fish.audio/plan
Fish Audio has a free tier. Paid plans start at $15/month. Source: https://fish.audio/plan
Fish Audio is best for video narrators, audiobook producers, and game studios who need line-by-line emotion control in synthetic speech instead of one flat voice setting across a whole script. Source: https://fish.audio/
Fish Audio is best for video narrators, audiobook producers, and game studios who need line-by-line emotion control in synthetic speech instead of one flat voice setting across a whole script. Skip it if vendor contradiction: the Free Tier plan card lists "Commercial use" while the homepage FAQ says the free plan is personal use only and the terms restrict unpaid use to personal, non-commercial use. Source: https://fish.audio/
The main Fish Audio features are S2.1 Pro real-time text-to-speech model, Inline emotion tags: [angry], [sad], [embarrassed], [emphasis], [whispering], [soft], [breathy], [excited], Special performance tags: [laughing], [chuckling], [clear throat], [sobbing], [crying loudly], [sighing], [panting], [groaning], [crowd laughing], [pause], [long pause], Voice cloning advertised at 15 seconds on the API section and as little as 10 seconds in the FAQ, and Speech-to-text with multispeaker labels, emotion tags, and natural language description. Source: https://fish.audio/
The closest Fish Audio alternatives on Citeware are Wave.video, VEED, Vidnoz AI, HeyGen, PixVerse. Full list: https://citeware.io/alternatives/fish-audio. Source: https://fish.audio/
Limitations called out on this listing: Vendor contradiction: the Free Tier plan card lists "Commercial use" while the homepage FAQ says the free plan is personal use only and the terms restrict unpaid use to personal, non-commercial use; Vendor contradiction: the annual toggle promotes "3 months off + anniversary 50% off when billed annually" but the billing FAQ says yearly billing saves 33% versus monthly; Vendor contradiction: cloning is advertised at 15 seconds in the API section and as little as 10 seconds in the FAQ; Vendor contradiction on languages: 30+ languages in the API section, 8 named languages in the FAQ, and 8 languages in the meta description; Unused minutes and credits do not roll over; monthly quotas reset at the start of each billing cycle; Free Tier caps each generation at 500 characters, too short for long-form narration; Credit math is indirect, with roughly 600 to 625 credits consumed per minute of generation; Professional voice slots are rationed by tier (1 on Plus, 5 on Pro, 15 on Max). Source: https://fish.audio/
Fish Audio is available on Web browser, REST API, On-premise deployment (Enterprise). Source: https://fish.audio/
Fish Audio integrates with Fish Audio REST API, Fish Audio SDKs for developers, Stripe (payment processing), Google sign-in via third-party account credentials, Voice Agent API for conversational applications, GitHub (fishaudio open-source repositories). Source: https://fish.audio/
Common Fish Audio use cases include Producing video voiceovers for YouTube, ads, and explainer content with scene-matched narration, Narrating audiobooks with chapter-level control aimed at ACX/Audible specs, Creating character voices and brand personas for games, animation, and interactive stories, Giving customer support bots and virtual agents low-latency conversational speech, Cloning a creator's own voice to produce narration without re-recording sessions, Transcribing multi-speaker recordings with emotion tags for editing or subtitling, Embedding speech synthesis into an app through the Fish Audio API. Source: https://fish.audio/
Fish Audio is made by Hanabi AI Inc. Address: 1111B Governors Ave STE 48109, Dover, DE 19904, United States. Official site: https://fish.audio/.
Fish Audio collects account and identity. Creating and managing your account or other user profiles. Fish Audio collects contact details. Corresponding with you; sending emails and other communications. Fish Audio collects payment information. Processing orders or other transactions; billing (shared with payment processor Stripe). Fish Audio collects product usage and analytics. Web analytics; improving the Services, including testing, research, internal analytics. Fish Audio collects device and ip data. Device/IP data shared with advertising and analytics partners; fraud protection, security and debugging. Fish Audio collects user content and uploads. User Content from input, file uploads or feedback used to provide the Services. Fish Audio collects cookies and tracking. Tracking Tools, Advertising and Opt-Out; collected automatically through Cookies. Fish Audio collects location. Geolocation data (IP-based location, browser time zone) shared with service, advertising and analytics partners. Fish Audio collects shared with third parties. Received from vendors, analytics providers, advertising partners and social networks; disclosed to service providers and business partners. Fish Audio collects advertising and remarketing. Marketing and selling the Services; showing you advertisements, including interest-based or online behavioral advertising. Fish Audio does not clearly say whether it collects used to train ai models. Not addressed; policy mentions improving the Services including research and development. Fish Audio does not clearly say whether it collects how long data is kept. Not specified in the provided text. Source: https://fish.audio/privacy
The Fish Audio pricing FAQ states each minute of generation costs roughly 600 to 625 credits. That places the 250,000-credit Plus plan at up to 200 minutes and the 2,000,000-credit Pro plan at up to 1,620 minutes, as listed on the Fish Audio plan page. Source: https://fish.audio/
No. Fish Audio states that monthly quotas reset at the beginning of each billing cycle and unused minutes do not roll over, so Fish Audio recommends spending the allocation inside the billing period. Source: https://fish.audio/
Yes. Fish Audio provides REST endpoints and SDKs for text-to-speech, speech-to-text, and voice cloning, and the Fish Audio pricing FAQ says premium subscribers get access to a flexible pay-as-you-go API with details in the developer documentation. Source: https://fish.audio/
A visitor can type into the demo box on the Fish Audio homepage, pick a voice such as Sarah or Adrian, add emotion tags, and generate audio before registering. Signing up for the Fish Audio Free Tier requires no credit card and unlocks the 8,000 monthly credits. Source: https://fish.audio/
Evidence: Homepage (Primary) · Verified Aug 21, 2026.
Looking for more options? Browse all Fish Audio alternatives.
Public threads about Fish Audio
Community observations