Audit before modeling
Attach your data in Ask and ask your research group to audit it. Explain the target, relevant features, and provenance you know. Inspect retained findings rather than treating successful execution as a clean bill of health.
Missing values, repeated rows, suspicious ranges, and provenance problems can change the question you should ask. Decide what each finding means for the study and record any corrective transformations as a new, traceable input.
Design an honest holdout
Tell your agents what the intended generalization is: new groups, later observations, or another supported boundary. Ask them to design and explain a suitable holdout using the supplied data.
Rows from the same experimental family may not be independent. If related observations cross partitions, a strong score can reflect overlap rather than useful prediction. Inspect partition membership and warnings before continuing.
- State what a future unseen example represents.
- Choose columns that capture shared experimental or source identity.
- Ask the agents to generate the split and inspect partition counts and group overlap.
- Keep the final test set out of model selection and record the intended evaluation protocol.
Compare supported baselines
Ask your research group to compare supported models under a common evaluation protocol. Supply the target, units, independence assumptions, and scientific context so the agents can plan the comparison.
Start with a simple supported baseline. Compare models under the same evaluation conditions. Inspect reported metrics together with sample counts, assumptions, warnings, and the generalization claim. The available scientific surface is determined by the installed packages.
Keep the test boundary meaningful
Do not repeatedly use final test results to choose the winning model. The workbench records evaluation protocols and test exposure, and some test outputs require an explicit reveal.
When a benchmark is refused, read the stated reason. Missing source declarations, unresolved independence, or invalid inputs should be corrected at their source. A passing software check does not replace review of the study design.