Methodology

How we rank, review, and compare AI models. Transparent, reproducible, and community-driven.

Our Review Process

1

Source Collection

We gather data from multiple independent sources:

  • Official Announcements: Company blogs, press releases, GitHub repos
  • Academic Papers: arXiv, conference proceedings
  • Standardized Leaderboards: MLPerch, LiveCodeBench, SWE-bench, LMSYS Chatbot Arena
  • Community Reports: Reddit (r/MachineLearning, r/LocalLLaMA), Hacker News, Discord/Slack
  • Financial Filings: SEC EDGAR, company filings
2

AI-Assisted Draft

Our AI (Allam 7B by SDAIA, reviewed by Hermes Agent) generates draft articles from collected sources. AI is used for research synthesis and draft generation only — not for factual claims.

3

7-Stage Verification

Each draft passes through 7 specialized subagents:

  1. Benchmark Analyst: Verifies benchmark claims against independent sources. Assigns confidence levels (HIGH/MEDIUM/LOW).
  2. Source & Fact-Check: Ensures every factual claim has a primary source link.
  3. Editorial Standards: Reduces hype language, ensures nuance, checks tone.
  4. Financial Analyst: Verifies financial claims, adds required disclaimers.
  5. Technical Depth: Verifies technical claims against documentation, removes fluff.
  6. Community & Trust: Ensures community perspective, checks for press-release feel.
  7. SEO Specialist: Ensures proper meta tags, internal linking, discoverability.
4

Human Review

Our editorial team reviews the final article for tone, accuracy, completeness, and editorial quality.

5

Publication

Articles are published only after passing all quality gates. Any article with >25% unverified benchmark claims is rejected.

Benchmark Scoring Methodology

Benchmark Categories

CategoryExamplesWeight
ReasoningMMLU, MMLU-Pro, GPQA, MATH25%
CodingHumanEval, LiveCodeBench, SWE-bench30%
LanguageMulBench, XCOPA, translation tasks15%
VisionMathVista, VQA, OCR15%
AgenticTerminal-Bench, tool use benchmarks15%

Confidence Levels

LevelRequirementColor
HIGH3+ independent sourcesGreen
MEDIUM1-2 independent sourcesOrange
LOWSelf-reported onlyRed

Ranking Criteria

Accuracy

Performance on standardized benchmarks. Weight: 40%

Efficiency

FLOPs per token, context window, memory footprint. Weight: 20%

Practicality

Hardware requirements, deployment ease, API pricing. Weight: 20%

Community Trust

Independent verification, community feedback, source quality. Weight: 20%

What Makes an Article "Verified"

Verification Checklist

  • ✓ All benchmark claims have at least 1 independent source
  • ✓ No "best ever" or "game-changing" without data support
  • ✓ Financial claims have SEC filings or official announcements
  • ✓ Technical claims verified against model cards/papers
  • ✓ Limitations and caveats included
  • ✓ Confidence levels assigned to all claims
  • ✓ Source links provided for every major claim
  • ✓ No fluff sections (empty "Future Directions")

Limitations

Our Limitations

We acknowledge that our methodology has limitations:

  • Not exhaustive: We cannot test every model on every benchmark.
  • Time-limited: AI releases move fast; some claims may become outdated.
  • Self-reported data: Some benchmark numbers are self-reported by companies and may not reflect real-world performance.
  • Community bias: Community feedback may reflect local lab preferences, not general consensus.

We strive for accuracy, but we are not perfect. Corrections are published within 24 hours when errors are found.

ZVHH Editorial Team

Last updated: June 29, 2026

Powered by Allam 7B (SDAIA)