AI Ark LLM Wiki

benchmarking

此標籤下有 5 條筆記。

  • 2026年9月29日

    SpaceXAI發布Grok 4.7!號稱「快一倍、半價」,實測為何跌破眼鏡?

    • raw-evidence
    • ai
    • model
    • benchmarking
    • pricing
    • token-economics
    • agent
    • harness
    • safety
    • reporting
  • 2026年9月29日

    GPT-6家族擴編!OpenAI推出兩款平價模型Sol與Luna:智力持平前代、API價格打5折

    • raw-evidence
    • ai
    • model
    • benchmarking
    • pricing
    • token-economics
    • agent
    • coding
    • reporting
  • 2026年9月29日

    Claude Sonnet 5.5有多強?程式功力大升級、文件、簡報製作也更成熟了

    • raw-evidence
    • ai
    • model
    • claude
    • benchmarking
    • pricing
    • token-economics
    • agent
    • coding
    • multimodal
    • reporting
  • 2026年9月29日

    Claude幫自己寫Code!Anthropic兩周讓App快3倍:如何用4撇步,讓AI自己找問題跟驗收?

    • raw-evidence
    • ai
    • ai-agent
    • claude
    • claude-code
    • performance
    • benchmarking
    • software-engineering
    • verification
    • governance
    • reporting
  • 2026年9月29日

    Anthropic推出Claude Opus 5.5!評測效能媲美Fable 5.1,運行成本大降40%成高CP值選項

    • raw-evidence
    • ai
    • model
    • benchmarking
    • pricing
    • token-economics
    • agent
    • coding
    • safety
    • reporting

Created with Quartz v5.0.0 © 2026

  • GitHub