Plicara Labs

My independent research practice: studies of AI evaluation, regex correctness, safety, and benchmark answer keys.

Plicara Labs is my independent AI research lab, where I investigate how AI systems behave through reproducible studies, benchmarks, and practical tools.

Regex evaluation

Can a generated regular expression pass its tests and still fail in use? I built regexbench and used it in a study of generated regular expressions, comparing test passing, reference-language equivalence, and denial-of-service vulnerability screening.

The study examines failures in generated patterns and in the benchmark’s own answer key. Its contribution is an inspectable account of what the metrics measure, where they disagree, and which limitations remain. Passing examples is not a general correctness or safety guarantee, and an automated equivalence check only answers the question within its supported semantics.

Read the study · Inspect the methods and results · Browse the experiment repository

Tools and further experiments

  • labloop, an agent-driven experiment loop that proposes a change, runs it time-boxed, and keeps it only if the metric improves.
  • regexbench, an evaluator for LLM-generated regular expressions covering semantic equivalence, correctness, and ReDoS safety.
  • Adventure Bench, a benchmark for interpreting an instruction as an action in a described scene, with explicit limits on what that task measures.

Read the research or follow along on GitHub.