Computational Text Analysis & NLP

Code & Text

I am a Digital Humanities Fellow at Princeton's Center for Digital Humanities and a research assistant in the digital humanities. Beyond that, I script, engineer, and solve — and create — problems for their own sake: I am a tinkerer and a digital do-it-yourself enthusiast.

I am currently at work on two research projects.

Fengliu and the ninth century

The first grows directly out of my doctoral dissertation, which reads the romantic narratives of the late Tang dynasty as sites of class tension. While working through these stories, my interest was caught by a tremendously elusive compound: fengliu 風流. Usually translated as "panache" or "sprezzatura", it behaved through the late Tang as a floating signifier — a word with no single referent.

Yet fengliu has a long and rich history reaching back to the Han dynasty, and across the centuries it took on radically opposed meanings. Having built, cleaned, and trained a corpus of texts spanning the Han through the early Tang, I am applying a range of text-analysis techniques specific to premodern Chinese, with the ultimate goal of tracing the semantic shifts of this compound. The project is nearing its final stages; I will release the code on my GitHub before long, once it has had a thorough tidying.

The material signature of woodblock prints

The second is a collaboration with the Department of East Asian Studies at Princeton. It sets out to recognise the material signature of woodblock prints, and through it to trace their movement across regions and centuries. My part is to build and label corpora of woodblock scans and to train models that detect cracks, cuts, and other marks left by the block itself.

Personal projects

I also write code away from research, for the pleasure of building my own software. Two projects I am fond of:

See more on GitHub →