Service overview · /eval/New to Medical Model Eval? See what it is, then book →
This page · item-set catalogAlready know the scenario? Pick an SKU and inquire—form at the bottom of this page

Eval dataset catalog · browse & inquire

Gate Catalog

Pick private sets by gate scenario; click Inquire to land on this page’s form.

This page is SKU selection & benchmark map, not a product manual. Service definition and booking live at /eval/. Inquiries sync to sales email, DingTalk, and CRM.

Reply within one business day · email · DingTalk · CRM · do not submit identifiable patient data · tel +86 15010400907

Which step do you want next?

Pick one of three. Opening a path prefills interest and drops you on the form.

Six eval dataset families · Inquire opens the form

Not sure which family? Submit a brief for our recommendation—or open Details for benchmark and paper breakdowns.

GB-ExamExam-style

Private difficulty-stratified sets

Aligned to MedQA / CMB / MedBench exam tracks. Fit for pretrain, SFT regression, and generation retests—public items contaminate easily; contracts need frozen private sets.

GB-ClinClinical dialogue

Multi-turn consult + physician rubric

Aligned to HealthBench / CMB-Clin. Release gates for consultation agents and clinical assistants cannot rely on MCQ totals alone.

GB-SafetySafety

Refusal · hallucination · jailbreak medical scenarios

High knowledge scores ≠ shippable. Safety tracks must clear separately before product and legal sign-off.

GB-VisionMultimodal

Imaging + report-understanding items

Derived from our authorized CT / WSI / checkup assets. A public CXR leaderboard does not buy you an authorization chain.

GB-AuditAudit

Contamination audit & difficulty calibration

An archivable report for procurement and financing diligence: leak checks, difficulty protocol, private-set freeze notes.

MG-LoopEval → data loop

Gate → failure buckets → patch pack

Evaluation finds gaps; buying data does not guarantee model lift. Gate outcomes map to scale-up, patch, or rollback-revision specs.

Already know which family?

Leave organization and scenario—we reply with an item-set briefing or Gate scheduling notes. SKU-named leads get priority scheduling.

Why buy private eval sets instead of grinding public boards?

Public benchmarks are a demand map; you pay for leak resistance, Chinese clinical rubrics, specialty difficulty, and auditable authorization—stable buys you cannot get from public leaderboards.

Buyer value

Contract-ready

Item-set version, scoring rules, isolation, and anti-contamination requirements can be frozen; gate outcomes are accept/reject—not a leaderboard screenshot.

Same stack

Same production line as our data

Same physician network, same data-card and license system. Gaps exposed by eval can loop into the next training-data pack specs.

No hype score

No unverifiable score theater

Public pages do not show unverifiable leaderboard scores. We deliver method, materials, and scheduling notes.

If you are benchmarking these public ecosystems

Open a public benchmark name for category details; on the detail page you can prefill an inquiry from that board.

You are benchmarking Common blockers Suggested inquire
MedBench / CMB exam track Open items contaminate easily; Chinese hard cases are scarce GB-Exam
HealthBench / CMB-Clin Open generation lacks physician rubrics GB-Clin
MedHallu / safety & ethics track High knowledge, low safety; hard hallucinations under-tested GB-Safety
CheXpert / multimodal track Authorization and report alignment cost high GB-Vision
Contamination / dynamic-eval diligence Unclear whether scores are memorization GB-Audit
Expand: delivery cadence and sample materials
01 · SubmitLeave scenario and intended SKU; get an inquiry ID
02 · AlignGate metrics, isolation, and scoring path
03 · MaterialsItem-set notes / Gate schedule / NDA when needed
04 · Loop backFailure types become next-pack specs

Sample:Gate report · Data-card fields · CT lung loop · Training dataset catalog

Submit eval brief

About 2 minutes. We reply with checkable materials: an item-set briefing or Medical Gate scheduling notes.

Tell us your gate goals

Leads that name an SKU, launch window, or need physician rubric / contamination audit are prioritized. After submit we sync sales email and DingTalk.

or only book Medical Model Eval Privacy

Operating & contracting entity: Beijing Langhui Technology Co., Ltd. · brand LumeSage. Campus names in office addresses (incl. PKU Science Park), do not imply university endorsement. This site is not a medical-device promotion and makes no clinical efficacy claims. Scale and license are per data card and contract. More: Notice · Privacy · IP · Compliance