data scientist II at skan.ai · bengaluru

i work on the parts of ML that sit close to the metal.

computer-use agents by day. triton kernels by night. i learn by building the thing and getting it wrong first.

cv ↗
now
leading a small team on a computer-use agent that drives real desktop apps for finance workflows.
nights
fused triton kernels in forge, and working my way into unsloth.
reading
how model weights leak the data they were trained on.

projects

things i built to find out. each clip is about seven seconds.

forge-kernels

gpu kernels · open source

fused triton kernels for fine-tuning, patched into a hugging face model with one call. every number comes from a committed result file, including the shapes where it loses.

  • 1.20× faster, 36% less memory on qwen2.5-0.5b
  • 660 tests
  • ★ 9

pact

latent reasoning · research

continuous thought without the recurrent unrolling. the model resynthesises its reasoning a stage at a time through a learnable query codebook, so training stops waiting on itself.

  • 4–9× faster training than coconut
  • same or better on gsm8k, prontoqa, prosqa
  • llama-3b / 8b

noctis

solana · hackathon

a fair value for tokenised us stocks during the 48 hours a week when no exchange or oracle says anything. published with an error bar, and cover priced off it.

  • live on mainnet data
  • program on devnet
  • stocklana, sept 2026

flyloom

connectome · simulation

the whole adult fruit-fly brain as a leaky integrate-and-fire model, fast enough on one cpu core that a lesion screen is a coffee break. you address neurons by name, not by 18-digit ids.

  • 138,639 neurons, 15m synapses
  • 36× faster than the matvec way
  • ships its own sanity checks

septa

on-device dictation

hold a key, talk, let go. a 0.6b model on your own machine turns the raw transcript into something you'd actually send. fillers gone, names and rupees left alone. no cloud.

  • fork of hex
  • mac, cleaner on windows

trading harness

autonomous research · paper only

the part that survives live markets isn't the agent debate, it's the evaluation machinery. llm agents write the journal and can veto a trade. they can't place one.

  • purged k-fold, next-bar fills
  • deterministic risk gate
  • india tax overlay on every pnl

also

work

all at skan.ai so far.

  1. mar 2026 — now

    data scientist II · computer-use agent

    leading a team on an agent that drives desktop GUIs across 12+ bfsi workflows. set-of-mark prompting and a thinker–grounder split got atomic steps to 95%.

  2. jul 2024 — apr 2026

    data scientist I · agentic process discovery

    next-activity prediction at 81% top-1 over 200+ classes, and the serving work that made it shippable: 2.9× throughput at 45% lower gpu cost.

  3. may — jul 2023

    intern · event extraction

    distilled bert, bart and pegasus into domain models 6× smaller at 98% of teacher f1.

  4. 2020 — 2024

    b.tech, electrical engineering · iit delhi

efficient training, kernels, agents, or a weird idea. happy to talk.

shaurya.m207@gmail.com

download the reel