๐Ÿ”Œ

multimodal

๐Ÿ”ŒPlugin

Orchestra-Research/AI-research-SKILLs

VibeIndex|
What it does
|

A collection of 7 multimodal AI research skills covering CLIP, Whisper, LLaVA, BLIP-2, SAM, Stable Diffusion, and AudioCraft โ€” part of Orchestra Research's 83 AI research engineering skills for coding agents.

Overview

A collection of 7 multimodal AI research skills from Orchestra Research, providing structured engineering instructions for coding agents to work with vision, speech, and generative models. Part of 83 AI research engineering skills designed to enable coding agents to conduct AI research experiments.

Key Features

  • CLIP โ€” OpenAI's vision-language model for zero-shot classification (320 lines, 25k+ GitHub stars)
  • Whisper โ€” Robust speech recognition supporting 99 languages (395 lines, 73k+ stars)
  • LLaVA โ€” Vision-language assistant for image chat at GPT-4V level (360 lines)
  • BLIP-2 โ€” Vision-language pre-training with frozen image encoders and LLMs
  • SAM โ€” Segment Anything Model for universal image segmentation
  • Stable Diffusion โ€” Text-to-image generation via HuggingFace Diffusers with SDXL and ControlNet (380 lines)
  • AudioCraft โ€” Meta's audio generation framework for music, sound effects, and audio

Who is this for?

AI researchers and ML engineers working on multimodal applications who want their coding agent to help with vision, speech, and generative model experiments. Ideal for teams building products that combine text, image, audio, and video understanding.

๐Ÿช

Part of

orchestra-research-ai-research-skills

Installation

Add marketplace in Claude Code:
/plugin marketplace add orchestra-research/AI-research-SKILLs
Step 2. Install plugin:
/plugin install multimodal@ai-research-skills
10,436
-
Last UpdatedJun 16, 2026