Skip to main content
Tag

benchmark

8 stories

Illustration for: Chinese AI models parrot state doctrine or refuse to answer
News

Chinese AI models parrot state doctrine or refuse to answer on sensitive topics Chinese AI models frequently toe the party line when asked politically sensitive questions, according to a study by Aleph Alpha. The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, giving it a commercial interest in distinguishing its models from Chinese competitors.

The Decoder3 min
Illustration for: Kling 4.0 Debuts 30s AI Video Model - Briefs Finance
Models & Research

Warning : Undefined variable $stocks in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : Undefined variable $funds in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 472 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d...

Kling AI14 min
Illustration for: Introducing LangSmith Fine-Tuning
Models & Research

Use the LangSmith Fine-Tuning CLI and skill - smithtune - to leverage agent trajectories stored in LangSmith to a fine-tuned model in one end-to-end workflow In addition to driving training, the smithtune CLI handles evaluating the fine-tuned model and uploads eval results to LangSmith for easy analysis We partnered with Fireworks & Baseten to create a seamless link between LangSmith...

LangChain Blog7 min
Illustration for: Humanity's Last Exam (Diamond) - Scale AI
Products & Tools

Challenging LLMs at the frontier of human knowledge HLE-Diamond, built with the Center for AI Safety (CAIS) , is a refined subset of Humanity’s Last Exam (HLE) resulting from a year-long process of cleaning and refinement with input from research communities.

Scale AI3 min
Illustration for: SWE-Bench Pro - Scale AI
Products & Tools

Evaluating challenging long-horizon software engineering tasks in public open source repositories SWE-Bench Pro is a benchmark designed to provide a rigorous and realistic evaluation of AI agents for software engineering.

Scale AI7 min