inference-testing
πPluginCloud-Byte-Consulting/plugins
Test and release models with evidence: model-level eval harnesses (synthetic ground truth, LLM-as-judge), GPU inference benchmarking (TTFT, tokens/sec, cost), release gates (explainability, fairness, adversarial robustness), rollout strategies (shadow/canary/bandit), 4-monitor drift stack, and GenAI red-teaming.
Part of
Cloud-Byte-Consulting/plugins
Installation
/plugin marketplace add Cloud-Byte-Consulting/plugins/plugin install inference-testing@cloud-byte-pluginsMore from this repository10
Assess an engineering org's platform and agentic readiness: ASDLC maturity scoring, platform ROI scorecard, org design, security/governance playbook, industry benchmarking, and IDP/ADP target architecture.
Engineering career coaching: senior/staff level-up behaviors, org signal reading, behavioral interview prep, AI-era positioning.
Thirty-five reusable workflows for model routing, personal productivity, code comprehension, knowledge systems, agent evaluation, consumer AI strategy, and Office documents.
Azure implementation arm of the platform suite: read-only estate assessment (Resource Graph, azqr, Governance Visualizer, aztfexport) feeding assessment increment I9, platform/landing-zone topology design on Radius and Azure Verified Modules, durable agentic operations on Dapr Workflow, four-layer IaC guardrail verification, workload onboarding by disposition, and the golden-path-as-an-API contract behind APIM with a platform MCP server on the roadmap.
Operate AI at work: model selection/routing, work-shape triage (chat/agent/team/nothing), agent output verification, cost/ownership/tool governance, and harness engineering (instructions, memory, handoffs).
Engineer the Agentic Developer Portal: CNCF platform-maturity benchmark with industry percentiles, MCP servers over platform APIs, agent identity (no standing secrets, ephemeral credentials), agent-consumable API contracts, golden paths as products, and fitness-function instrumentation for the maturity roadmap.
Make research data collection trustworthy without gatekeeping: ODCS data contracts with CI enforcement, dataset QoS/SLOs, data-product reviews (DAUTNIVS), right-sized governance with steward roles and certification, agent-consumable dataset catalogs, and lakehouse storage architecture for training data.
Ship production LLM apps: eval engineering (tracing, RAG metrics, LLM-as-judge, guardrails) and deployment (containers, cloud rollout, vector-store tuning).
Azure DevOps CI/CD and full-SDLC assessment: CLI-first discovery (az + azure-devops extension), Wiki-ready Mermaid documentation, governed pipeline architecture with 30/60/90 roadmaps, adversarial PR review (Advocate / Skeptic / Judge), and chapter-indexed domain reasoning across twenty-five comprehensive engineering book guides.
Run model creation and training as a product: training pipeline architecture (Argo/Kubeflow, trigger taxonomy), MLflow experiment/registry standards, distributed-training topology selection (Ray/Dask/Spark), Ray on K8s operations, notebook-to-production golden path, and fine-tuning strategy (prompt vs RAG vs PEFT).