galaxy
StartupsFundersInstitutionsPeopleNewsMap
Admin
Startups
company

Vals AI

valsai.click →

86profile quality

Vals AI builds independent, proprietary AI benchmarks and evaluation infrastructure to measure model performance in finance, law, coding, and public sectors.

ai
Business Model Canvas · v7

Value proposition

"The independent evaluator of artificial intelligence" — Vals builds proprietary, privately held benchmarks and evaluation infrastructure that measure whether AI models can perform economically valuable work in finance, law, software engineering, and healthcare, providing a credible, third-party signal of model capability and reliability that labs and enterprises lack.

Where it wins

  • Private test sets prevent training leakage: Unlike open-source benchmarks absorbed into pre-training corpora, Vals keeps its test suites private to preserve signal integrity and prevent reward hacking [1].
  • Agentic and long-horizon evaluation: Vals measures real-world utility through multi-modal, tool-use, and long-horizon tasks (e.g., 12- to 30-hour autonomous R&D budgets) rather than static exam-style questions [2][1].
  • Industry-standard adoption: Its benchmarks (e.g., Finance Agent, Legal Research Bench) are rapidly becoming industry standards, with the world's leading AI labs (Anthropic, DeepSeek, Google) relying on Vals to test models before release [1][3].
  • Public good and government rigor: Vals brings independent, third-party rigor to high-stakes public domains (e.g., SNAP benefits, cybersecurity) with rubrics validated by domain experts, addressing the trust gap for government adoption [1][3].

Credibility: Vals benchmarks are covered by the New York Times, Washington Post, Wall Street Journal, and Bloomberg, and its infrastructure (Valkyrie) is built for reproducible, large-scale evaluation [1][3].

123

Business model

  • Proprietary Benchmark Creation: Vals builds high-quality, privately held test sets in collaboration with domain experts, preventing test-set leakage and maintaining evaluation integrity [1].
  • Infrastructure-Led Scale: It uses its own distributed system (Valkyrie) and model library to run evaluations reproducibly and at scale across multiple labs, creating a scalable evaluation engine [1].
  • Standard-Setting via Leaderboards: By publishing live, frequently updated leaderboards (e.g., Vals Index, RSI Index), Vals establishes its benchmarks as industry standards, driving demand for its evaluation services [2][3].
  • Trust as a Service: Vals monetizes the lack of independent measurement in the AI market by providing a credible, third-party signal of model performance and risk, similar to rating agencies in finance [1][3].
  • Multi-Domain Expansion: It expands its value proposition by adding new benchmarks across economically valuable and scientifically important domains (finance, law, coding, public benefits) [1].
123

Competitive landscape

  • Open-Source Benchmarks (e.g., MMLU, SWE-bench): Widely used but suffer from test-set leakage and reward hacking; Vals counters with private test sets and rigorous, agentic evaluation [1].
  • Lab-Run Evaluations: AI labs often evaluate their own models, leading to potential bias and cherry-picking; Vals provides independent, third-party validation [1].
  • Commercial AI Rating Agencies: Emerging firms offering AI ratings, but Vals differentiates through deep domain expertise, proprietary infrastructure, and government/public sector focus [1][3].
  • Academic Benchmarks: Often contrived and exam-style, lacking real-world utility; Vals focuses on economically valuable, multi-modal, and long-horizon tasks [1].
  • Differentiators: Vals' private test sets, agentic evaluation infrastructure (Valkyrie), industry-standard adoption by leading labs, and trusted public sector presence create a defensible moat [1][3].
123

Market pains

  • Lack of Independent Measurement: The AI market lacks credible, third-party evaluation institutions, making it difficult for labs to demonstrate progress and for enterprises to quantify ROI [1].
  • Test-Set Leakage and Reward Hacking: Open-source benchmarks are often absorbed into pre-training corpora, invalidating results and encouraging models to game evaluations rather than improve [1].
  • Trust Deficit in High-Stakes Domains: Government agencies and enterprises struggle to trust AI with critical decisions (e.g., benefits, legal, cybersecurity) without independent evidence of reliability [3].
  • Rapid Pace Outpacing Evaluation: The speed of AI development exceeds the community's ability to construct new, meaningful benchmarks, leading to models summiting old tests quickly [1].
  • Cherry-Picked Reporting: Labs often report evaluations alongside cherry-picked examples and tuned regimens, undermining transparency and fairness [1].
123

Strategic implications

Vals has successfully positioned itself as the 'Moody's of AI,' solving a critical trust and measurement gap in a trillion-dollar market. Its private test sets and agentic evaluation infrastructure create a defensible moat against test-set leakage and lab-run biases. The main risk is the potential for AI labs to develop internal evaluation capabilities or partner with alternative providers, though Vals' industry-standard status and government trust mitigate this. The opportunity lies in expanding into new high-stakes domains (e.g., healthcare, climate) and deepening government partnerships, which could lead to regulatory mandates for independent AI evaluation. The next signal to watch is whether Vals' benchmarks are adopted as de facto standards for AI procurement or regulation, which would solidify its market position and pricing power.

123

Improvement suggestions

Vals should develop a transparent, auditable methodology for its private test sets to further enhance trust among skeptical enterprises and regulators, potentially through third-party audits or open validation sets. It should expand its government partnerships by offering tailored, compliance-focused evaluation packages that align with emerging AI regulations, creating a sticky, high-value revenue stream. Vals could also monetize its infrastructure by offering a 'Vals-as-a-Service' platform for enterprises to run custom benchmarks, leveraging its existing Valkygie system. Finally, Vals should invest in building a community of domain experts and researchers to co-create benchmarks, fostering innovation and expanding its benchmark library faster than competitors.

123
Sources
  1. https://www.vals.ai/about import · fetched Sep 2, 2026
  2. https://valsai.click/ import · fetched Sep 2, 2026
  3. https://www.vals.ai/gov import · fetched Sep 2, 2026

Overview

Country
Not verified
City
Omitted: No headquarters city is stated in the provided documents.
Stage
Seed
Categories
ai
Profile completeness
5 of 6 fields
Last researched
Aug 15, 2026
Quality score
86/100