langsmith-dataset
๐ฏSkillfrom langchain-ai/skills-benchmarks
Part of Claude Code Skill Benchmarks by LangChain, a framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Provides Docker-sandboxed task execution with configurable treatments, repetition counts, and parallel workers for systematic evaluation.
Same repository
langchain-ai/skills-benchmarks(21 items)
Installation
npx vibeindex add langchain-ai/skills-benchmarks --skill langsmith-datasetnpx skills add langchain-ai/skills-benchmarks --skill langsmith-dataset~/.claude/skills/langsmith-dataset/SKILL.mdSKILL.md
More from this repository10
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with treatment-based test variations.
A benchmark suite by LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and repetitions.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It supports multiple treatments, repetitions, and parallel test execution with Docker-sandboxed validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and automated validation.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. Supports multiple treatments, repetitions, and parallel test execution in Docker-sandboxed environments.
A benchmark skill from LangChain that measures how skill documentation design affects Claude Code's adherence to recommended patterns, with support for multiple treatments and configurable model selection.
Part of LangChain's skill benchmarks project that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed test tasks with configurable treatments.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns, using Docker-sandboxed tasks with configurable treatments and validation scripts.
A benchmarking framework that measures how skill documentation design affects Claude Code's adherence to recommended patterns. It runs tasks in Docker-sandboxed environments with configurable treatments and repetitions to evaluate skill effectiveness.
A benchmarking framework from LangChain that measures how skill documentation design affects Claude Code adherence to recommended patterns, using configurable treatments, Docker-sandboxed execution, and automated validation.