Independent scholar — agency and selfhood in artificial intelligence

For fifteen years I built and ran the systems a university used to find out whether its students were actually learning what it claimed to teach them.

That work is usually called learning assessment. It is more accurately described as evaluation design under adversarial conditions: the people supplying your data have no particular incentive to supply it, the outcome you are measuring is sometimes not directly observable, the measures available are proxies of uneven quality, and the results will be read by an external body with the power to sanction the institution. Everything I know about measuring hard-to-measure things, I learned there.

Designing instruments

I designed instruments as well as the analyses: surveys, exit surveys, focus group protocols, and rubrics. The largest was a study of roughly 1,800 undergraduates establishing how students understood a required writing sequence — what they believed they had learned, and which skills they expected to keep using afterward. I have built instruments in Qualtrics, run outcomes collection through Canvas, and designed faculty-facing reporting forms customized per college so that a single submission could serve multiple institutional purposes at once.

The recurring design problem was never statistical. It was that the instrument had to survive contact with people who potentially did not want to fill it out, and who came armed with misconceptions about what learning assessment is for.

Operationalizing outcomes

Most of the work was turning vague institutional aspirations into things that could actually be measured, and then getting hundreds of separate programs to agree on the result. I wrote guidance the campus used for stating program learning outcomes in a form that could be assessed, and drove a campus-wide project to bring every degree program’s outcomes into a standard format and onto the public catalog.

My work focused on two separate processes: program assessment asks whether students leaving with a degree achieved that program’s outcomes. General education assessment asks whether learning happened in individual courses carrying general education credit, against university-wide outcomes. These are different processes with different data sources, different reporting chains, and different failure modes.

Disaggregation, and finding where a system fails a subgroup

The analysis I care most about is disaggregation: not “how did students do,” but which students did worse, and at which specific point in the curriculum the gap opens. An aggregate pass rate conceals the thing you need to know. The useful output is a location — this course, this outcome, this population — because a located gap is actionable and an aggregate number is not.

This is the part of my background that transfers most directly to evaluating AI systems. It is the same operation as subgroup analysis in a fairness evaluation: the headline metric is the least informative number in the report.

Meta-evaluation

I built and used a rubric for reviewing other programs’ assessment plans: whether the measures actually bore on the stated outcomes, whether the sampling supported the claims being made, whether the reported action plans followed from the findings. Much of my consulting work was telling people, constructively, that their evidence did not support their conclusion.

A related and underrated part of the job: coaching the people who use data on what a given dataset can and cannot support. The most common institutional failure I saw was not bad data. It was good data asked to answer a question it was never built to answer.

Building infrastructure, and knowing when to stop buying it

I ran the campus’s assessment records, reporting cycles, and compliance tracking for more than 200 academic programs and every general education course. When a commercial assessment platform proved to be a poor fit, I replaced it with purpose-built forms, structured shared storage, and spreadsheet-based tracking — and produced better accreditation reporting at lower cost. A second commercial system, used seriously by only a handful of units, was replaced by a workflow tool that got institution-wide compliance because it fit how people actually worked.

The general lesson, which I would carry into any evaluation role: adoption is a design constraint, not an afterthought. A rigorous system nobody uses is worse than a coarse one everybody does, because the first produces the illusion of measurement.

Reporting to people with authority over you

I prepared the institution’s reporting on student learning for its regional accreditor and for the state system, including the evidence used in a reaccreditation cycle. Writing for a body that can sanction you is a specific discipline: claims have to be supported by evidence you can produce on request, gaps have to be named before the reviewer names them, and the difference between what you measured and what you wish you had measured has to be stated in your own words first.

What I do with it now

I write about artificial intelligence, and I am interested in evaluation as it applies to model behavior: how you build an instrument for a capacity you cannot observe directly, how you tell a real signal from an artifact of the measure, and how you find the subgroup the aggregate is hiding.

Artifacts from this work — instruments, rubrics, reports, and analyses — are available on request.