Calibrating Generated Items Before Pretesting with Predicted Priors

Predicts Rasch item difficulty from item features (text embeddings from any model, template family, content metadata) by ridge regression (Hoerl and Kennard, 1970, ) with honest out-of-sample uncertainty, uses the predictions as robust informative priors to calibrate new items from small pretest samples, in the spirit of using collateral information in calibration (Mislevy, Sheehan and Wingersky, 1993, ), plans the pretest sample size needed for a target precision, and diagnoses item families whose predictions cannot be trusted.


coldstart

Calibrate AI-generated items before you have the pretest seats to do it the old way.

Automatic item generation and LLM-assisted writing produce items faster than pretesting can calibrate them. coldstart predicts each new item's difficulty from its features, uses the prediction as a robust prior, finds the families where prediction fails, and tells you how many responses each item needs.

library(coldstart)

sim  <- cs_simulate(seed = 1)                     # legacy + new items with features
it   <- sim$items; tr <- it$set == "train"

pr   <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr])
pred <- predict(pr, sim$features[!tr, ], it$family[!tr])   # mean + honest SD

cs_plan(pred, target_sd = 0.3)                    # responses needed per item

resp <- cs_responses(sim, n_per_item = 25)        # or your pretest data
cal  <- cs_calibrate(resp, pred)                  # t-prior Bayes vs baseline
chk  <- cs_check(cal, setNames(it$family[!tr], it$item[!tr]))
cal  <- cs_calibrate(resp, cs_distrust(pred, chk))  # drop priors that failed

Installation

From CRAN (once released):

install.packages("coldstart")

Development version from GitHub:

install.packages("pak")
pak::pak("edidatasolutions/coldstart")

Features can be anything numeric: embeddings from any text model, cognitive attribute codes, content metadata. The package does not call a model itself.

Design choices

  • Honest predictive SD, two kinds. Out-of-fold RMSE for families seen in training, and leave-one-family-out RMSE for new templates.
  • Robust prior. The default Student-t (df = 4) lets the data override a bad prediction instead of being dragged toward it.
  • Family-level trust check. A chi-square test of prior-data conflict per family; cs_distrust() withdraws priors from failing families.

Validation (known truth, 5 replications, inst/validation/known_truth.R)

Predictive SDs are honest: stated 0.54 vs actual RMSE 0.53 (seen families), 0.70 vs 0.64 (unseen family); 90% intervals cover 90.8% and 90.4%.

RMSE of difficulty, seen families:

responses per item baseline predicted prior
15 0.72 0.40
25 0.53 0.35
50 0.36 0.30
100 0.27 0.23

With the prior, 25 responses do what 50 do without it.

Drifted ("rogue") template family (+1.2 logits vs its history): the prior alone hurts (0.66 vs 0.61 at n = 25). cs_check flags the family in 80% of replications at n = 25 and 100% at n ≥ 50, with 0.2 false family flags per replication. After cs_distrust(), RMSE is back at baseline (0.57 at n = 25).

Planner: a target posterior SD of 0.30 needs a median of 40 responses per item with the prior vs 57 without; achieved SD 0.302, RMSE 0.300.

Status

Done: cs_simulate, cs_responses, cs_predictor (+predict), cs_calibrate, cs_plan, cs_check, cs_distrust. Next: 2PL (discrimination priors), sequential updating as responses arrive, ability uncertainty for pretest examinees (currently treated as known from operational scoring), and non-linear predictors.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("coldstart")

0.1.0 by Daniel Edi, 21 hours ago


https://github.com/edidatasolutions/coldstart, https://edidatasolutions.github.io/coldstart/


Report a bug at https://github.com/edidatasolutions/coldstart/issues


Browse source code at https://github.com/cran/coldstart


Authors: Daniel Edi [aut, cre, cph] (ORCID:


Documentation:   PDF Manual  


MIT + file LICENSE license


Imports stats

Suggests knitr, markdown


See at CRAN