Fish Audio Raises $50M Seed to Build AI Voice Models for Creators and Enterprises
Fish Audio secures $50M seed funding led by Coreline Ventures to expand its AI voice models, with 8 million users and $21M annual recurring revenue—positioning itself as an ethical leader in the growing AI voice market.
Fish Audio Secures Major Funding to Expand AI Voice Technology
What happened
Palo Alto-based startup Fish Audio has successfully raised $50 million in a seed round, marking a significant milestone in the rapidly evolving AI voice technology landscape. The funding was led by Coreline Ventures and Capital Today, with additional participation from an impressive roster of investors including 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.^1
The company has already demonstrated substantial traction in the market. Fish Audio's library contains more than 15,000 natural language controls for AI voice models, providing creators with unprecedented flexibility in generating and customizing AI voices.^2 Since launching last year, the startup has attracted more than 8 million users who are leveraging its open-source or hosted versions of its models.^3
The company's financial performance is equally impressive, currently generating annual recurring revenue of $21 million.^4 This strong revenue stream underscores market demand for AI voice solutions and positions Fish Audio as a serious contender in an increasingly crowded field. The company's technical credibility is further evidenced by its Fish Speech repository on GitHub, which has garnered more than 31,000 stars from the developer community.^5
In terms of product development, Fish Audio has been remarkably active, launching five models in the last year: four speech generation models and one speech-to-text model.^6 Notably, three of its speech generation models have been open-sourced, while the latest S2.1 Pro model is available exclusively through its paid API.^7 This hybrid approach allows the company to build community trust while monetizing its most advanced capabilities.
Perhaps most importantly for users concerned about content rights and privacy, Fish Audio's CEO Rissa Cao has announced a significant new feature: creators can now have their voices taken off the platform in less than 3 minutes using an automated take-down process.^8 This rapid response capability addresses one of the most pressing concerns in the AI voice space—consent and control over one's digital identity.
The speech generation market is highly competitive, with established players like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp all vying for market share.^9 Fish Audio's ability to secure such significant funding from top-tier investors suggests it has a compelling value proposition that resonates with both creators and enterprises.
Why it matters
The $50 million seed round for Fish Audio represents more than just another funding announcement in the AI sector—it signals a fundamental shift in how creators and enterprises approach voice technology. The involvement of prominent investors like Coreline Ventures, whose partner Oskue Honda emphasized that "consent, transparency, and attribution must be built into AI voice platforms,"^10 indicates that this round was driven by more than just growth metrics. These investors are betting on a future where ethical considerations around AI-generated voices become central to product development and market adoption.
The company's approach of open-sourcing multiple models while keeping its most advanced offerings behind a paid API is particularly noteworthy in the current AI landscape. This strategy mirrors successful patterns seen in other sectors—building community trust through open-source contributions while creating sustainable revenue streams from enterprise-grade solutions. The 31,000+ GitHub stars for the Fish Speech repository^11 demonstrate that this approach has resonated strongly with developers and technical users.
The rapid automated take-down process, allowing creators to remove their voices in less than 3 minutes, addresses a critical pain point in the AI voice industry. As voice cloning technology becomes more accessible, concerns about consent and unauthorized use have grown exponentially. Fish Audio's solution provides a practical mechanism for maintaining control over one's digital identity, potentially setting a new standard that other players in the space will need to match or exceed.
The company's impressive user base of 8 million people using its models^12 suggests that Fish Audio has already achieved significant product-market fit. This is particularly remarkable given that it launched just last year, indicating rapid adoption and strong value proposition in a competitive market.
For enterprises specifically, the combination of advanced AI voice capabilities with enterprise-grade security and compliance features (as evidenced by the automated take-down process) makes Fish Audio an attractive option for organizations looking to implement voice AI solutions without compromising on ethical considerations or user consent.
What to watch
Several key developments will be important to monitor as Fish Audio continues its growth trajectory:
Product Evolution: With five models launched in the last year, including three open-sourced speech generation models and one speech-to-text model, Fish Audio has demonstrated an aggressive product development pace.^13 Investors will want to see how the company maintains this momentum while managing technical debt and ensuring quality improvements with each release. The S2.1 Pro model's exclusive availability through the paid API^14 suggests a tiered product strategy that could be refined over time.
Market Position: With competitors like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp all operating in the same space,^15 Fish Audio will need to differentiate itself beyond its ethical commitments. Key differentiators to watch include model quality, pricing strategy, API performance, and developer ecosystem building.
Ethical Leadership: The company's emphasis on consent, transparency, and attribution—championed by Coreline Ventures partner Oskue Honda^16—could become a significant competitive advantage as regulatory scrutiny around AI-generated content increases. Fish Audio's position as an ethical leader in the space could attract both users concerned about privacy and enterprises requiring compliance with emerging regulations.
Community Engagement: The strong GitHub presence (31,000+ stars for Fish Speech)^17 indicates an engaged developer community. How well Fish Audio nurtures this community through documentation, support, and open-source contributions will be crucial to its long-term success.
By the numbers
Source snapshot
