About Me

I came to evaluation sideways.

My PhD at the University of Georgia, with Sheng Li, was about knowledge — specifically, how to get structured knowledge that language models were never trained on into the models anyway. Domain-adapted encoders for agriculture and clinical text, knowledge graphs built out of raw documents. Along the way I spent time at Adobe Research and Thomson Reuters. It was good work, and the thing I kept running into was not the modeling. It was that I could never fully convince myself the numbers meant what everyone assumed they meant.

That question is now most of my job. At NBME I work on automated scoring of how doctors communicate with patients — building the systems, and then trying very hard to break my own evidence that they work.

I like it here because the scores matter. They feed into how physicians get assessed, so there are experts reading over my shoulder and real pressure to get it right. That turns out to be a very good way to find out whether a method actually works or just looked good on a slide.

Outside of work I play the setar and wander around cities with a camera, usually at the wrong shutter speed on purpose.

Always happy to talk shop — the links up top all reach me.