BrandonBaek

AI/ML researcher · Educator · Community builder

Retrace
Brandon Baek in a white shirt and tie

About me

I’m a Korean-American student and aspiring AI/ML researcher interested in decision systems and operations research.

La Crescenta, CA Korean-American

Explore

Research

연구

Three first-author papers. Solo work on AISTATS; team research on ExactPHQ and RECHECK-ED.

AISTATS 2027

Do Adaptive Neural Networks Learn Computational Depth or Surface Difficulty?

Main-track submission · Under review

Solo author

Brandon Baek

17-page main-track manuscript · Public overview

A right answer is not always a usable exit.

My solo-authored methodological paper asks whether adaptive neural networks respond to computational depth or surface difficulty. Its central distinction is between an early prediction that happens to be correct and a stopping decision supported by the information the controller has.

17pages · main-track submission
Soloauthored by Brandon Baek

Available information comes first.

An evaluator with the answer can recognize a lucky early guess. A deployable stopping controller cannot use that future knowledge. The research audits that information boundary.

What can the exit rule actually see?

A conceptual illustration, not an experimental result.

A?0
B?1

Matching observation histories require the same stopping decision, even if the eventual answers differ.

Controlled reasoning tasks

The manuscript investigates adaptive neural computation through controlled reasoning tasks. The public page explains the question and its information constraints while the full study is under review.

Prediction

What answer does the model propose?

Observation

What information has arrived so far?

Stopping

Can the controller justify ending computation now?

Submission status

Submitted to the AISTATS 2027 main track. The manuscript and detailed numerical results stay private during review; no acceptance is implied.

The concept illustration above explains the research question; it is not a published finding or a reproduction of a private paper figure.

Compute savings need an attainable stopping rule.

The broader motivation is to separate savings that look possible with hindsight from savings a model can choose in operation. This is a methodological question about evidence available at decision time, rather than another applied classification project.

ExactPHQ

ExactPHQ: Certified Stopping with Learned PHQ-9 Ordering

IEEE ICDM 2026 · Teen Research Track · Camera-ready complete

Shenyang, China

First author · Collaborative research

Brandon Baek · Tevyn Yong · Seonho Kim

Learning orders. Bounds decide.

ExactPHQ asks a narrow question: can fewer PHQ-9 items determine the same binary threshold class as all nine? A learned policy chooses what to ask next. A deterministic certificate alone decides when to stop.

25.22%fewer questions · primary external replay
262,144fixed patterns verified · zero class errors

These are retrospective question counts and exact fixed-response score classification, not clinical equivalence.

The stopping certificate

For answered items, L is the observed sum. With m unanswered items scored 0–3, U = L + 3m. At T = 10, stop positive if L ≥ T; stop negative if U < T. Otherwise another response is necessary. Model confidence cannot narrow these bounds.

Follow the paper’s worked example

A fixed response vector from Table 2. This illustrates score bounds, not a screening tool.

I9 ?I1 ?I2 ?I3 ?I4 ?4 more
[0, 27]The threshold is unresolved.

Item 9 is captured first. That protocol does not establish suicide-risk assessment or a safety intervention.

Freeze, replay, replicate.

We developed policies on adult NHANES 2017–2018 records, locked them, and replayed complete responses on two different cycles. The newer cycle is the primary external evaluation; the earlier cycle adds replication.

5,068Development · 2017–2018
5,455Evaluation · 2021–2023
5,134Replication · 2015–2016

Empirical dynamic programming plans future question cost. Secondary policies use a categorical MADE density or a compact latent-class mixture. Learning changes the order, never the stopping rule. Later searches are exploratory because they reuse already examined cohorts.

A smaller learning gain, honestly shown.

Where the question savings come from

Survey-weighted mean questions. All four policies preserve the binary class for fixed responses.

Full PHQ-99.000
Optimal fixed order6.741
Empirical DP · primary6.731
MADE DP · exploratory6.650

n = 5,455 adults · Mean questions out of nine

Most savings come from certified stopping and a good fixed order. In the primary cohort, empirical DP saves only 0.010 questions beyond optimal fixed ordering. Exploratory MADE saves 0.090 beyond fixed and 0.041 beyond an independent-item model.

All 4⁹ fixed patterns had zero class errors in independent exhaustive traces. Preserving five severity bands instead required 8.394 questions in the newer cohort.

What the guarantee covers

The guarantee is exact binary score emulation for a fixed response vector. It does not say a person will give the same responses after reordering. A prospective study would need to measure response consistency, time, dropout and distress, with independent human escalation rules.

  • No claim of diagnosis, time savings, or prospective clinical equivalence.
  • The binary threshold guarantee does not preserve an exact total or complete severity assessment.
  • The work was coauthored by Brandon Baek, Tevyn Yong and Seonho Kim.

RECHECK-ED

Learning Physiological Visibility to Time Repeated Emergency Nursing Reassessments

IEEE BigData 2026 · High School Symposium · Decision pending

Phoenix, Arizona

First author · Collaborative research

Brandon Baek · Seonho Kim · Tevyn Yong

A change can disappear before the next check.

RECHECK-ED studies when to place one already permitted flexible nursing reassessment. It combines recent waveform support with the probability that a visible change ends within 20 minutes, while keeping required checks fixed.

+9.0 ppcapture versus fixed +40 · equal check budget
11 fewerendpoint-negative prompts versus raw monitoring

A retrospective information-capture study, not a trial of outcomes, nurse workload, or safety.

Timing without moving the protected checks

Two patient-separated, five-member gradient-boosted ensembles estimate waveform-rule support q and 20-minute disappearance e from past information only. Their product q × e forms a short-horizon urgency heuristic. Missing waveforms cannot reassure the controller: an active raw-monitor event triggers a check.

One flexible check. A protected deadline.

Select an offset to inspect the frozen holdout action trace. These are block counts, not a patient simulation.

+120

556 / 648 blocks selected +60 min

Minutes within each complete two-hour block. Initial and +120 checks stay fixed; required and urgent checks are never movable. This repeats in later complete blocks.

Patients separated. Method frozen.

The method was developed on 450 previously opened MC-MED waveform episodes across 412 patients. Model files, thresholds, code and waveform inventory were frozen before a one-time evaluation on 600 new episodes from 567 different patients.

600selected holdout episodes
402met the waveform rule
200qualified and not already chart-represented

The other 202 waveform-qualified episodes were already chart-represented. Thirty-one of the 600 episodes had no complete waveform overlap. The case-enriched, single-center sample does not estimate general ED event prevalence.

Capture and selectivity are different gains.

Same check budget. Different information capture.

200 eligible endpoints; 1,671 attempted checks per policy.

The model preserves the raw monitor’s 70 captures while reducing endpoint-negative prompts from 67 to 56.

The model’s +9.0 percentage-point difference from fixed timing has a patient-cluster 95% interval of 5.08–13.07. Capture is overlap with the simulated reassessment, not a measured clinical response.

The raw-monitor trigger captured the same 70 episodes as the model. Learning’s incremental result was 56 endpoint-negative triggered actions instead of 67—not fewer total attempted checks. Retrained no-waveform and process-only models also captured 70, with 62 and 60 endpoint-negative prompts.

0.950waveform-support AUROC
0.75320-minute expiry AUROC
0protected-check violations

A research prototype, governed by nurses.

The result supports monitor-responsive timing and a smaller set of endpoint-negative prompts under a fixed experimental budget. It does not establish that nurses missed events, that prompts were unnecessary, or that waiting would be safe. Required, urgent, post-medication and post-procedure checks cannot be moved.

  • PPG pulse quality cannot confirm arterial oxygen saturation.
  • The replay leaves physiology, treatment and historical documentation unchanged.
  • Prospective work needs nursing review, local permissions, override rules and measured patient outcomes.

Projects

프로젝트

Models, websites, and useful experiments.

Full record

Choose a collection to browse. Search is here if you need something specific.

Topic

Loading archive…

What I’m working on next
01GovGuideNavigating local government, with clarity.In development

A government-navigation project focused on making local government processes easier to understand.

View project →
02Deep learning courseMachine learning and deep learning, built on basic calculus.In development

A machine learning and deep learning course built around basic calculus and core concepts. Not yet published.

View project →
GitHub

Community

함께

AI ethics, education, music, and community service.

01

HumanityOverAI

Founder & leader

I founded and lead a ten-person student organization making AI ethics education accessible through curriculum, public events, community challenges, and research.

  • Designed a 56-question AI perception survey with 1,651+ responses across 27 countries; a formal report is in development.
  • Presented at an expo and educated 120+ adults about AI’s impact.
  • Built team workflows and AI ethics lessons, and organized an awareness challenge with about 22 participants.
02

Korean American Music Academy

Regional student president

Since 2025, I’ve served as regional student president at KAMA, a Korean-American music nonprofit supporting community causes, including bone marrow donor searches and musicians with disabilities.

  • Led 20 students across events, rehearsals, and community performances.
  • Planned four major events, including concerts and banquets, and conducted three concerts.
  • Raised funds for bone marrow transplant awareness and musicians with disabilities.
  • Expected to take on the full student presidency in 2027.
03

Homemade Delights

Co-founder & head of marketing

I co-founded and co-lead a nine-person organization connecting mental health awareness with community service. I lead marketing, research-backed writing, outreach, and events.

  • Helped raise $2,500+ for local charities, including $600 for mental health awareness, and supported sales of 500+ baked goods.
  • Supported a 55-question survey with 2,300+ responses across 30+ countries; helped the team win a $750 survey contest.
  • Wrote about half of roughly 300 research-backed articles in one year and organized holiday baking and donations for people in need.
04

Machine learning & deep learning teaching

Teacher · Course developer

My teaching now focuses on machine learning and deep learning. I’m developing a course that builds on basic calculus to explain the core concepts.

  • Previously tutored 15+ students aged thirteen and up in computer science and AI.
  • The new machine learning and deep learning course is in development; it is not yet published.
05

AI Research Team

Team leader

I lead a separate five-member student research team exploring AI and data science, including student well-being prediction, online creator feedback, and housing-price modeling.

  • Set project directions and coordinate research and development.
  • Mentor teammates and review methods and results.
  • Oversee open-source documentation and reproducible workflows.
06

Student Self Defense Advocates

Team member

I work with a five-person student advocacy team on student independence and safety through research, outreach, and public events.

  • Contribute to research, outreach, and events with a five-person team.
  • Helped gather 900+ survey responses across 20+ countries.
  • Supported approximately $1,250 in fundraising for student self-defense awareness.
07

Next Generation Advocates

Volunteer

Volunteer at Los Angeles community and Korean American heritage events, and contribute to projects around mental health infrastructure and migrant heritage culture.

  • Volunteered at Los Angeles community and Korean-American heritage events.
  • Recorded 20 volunteer hours at the 105th Commemorative Ceremony of the Provisional Government of the Republic of Korea.

Service recognition

The certificates behind the work, including my 2024 Presidential Volunteer Service Awards.

About me

소개

Korean-American. La Crescenta, California.

I’m Brandon, a Korean-American student in the LA area. I research and teach machine learning, and lead student projects in AI ethics and community service.

School
Crescenta Valley High School · Class of 2028
Languages
English / 한국어
Research
AI/ML · Decision systems · Operations research
My research
Brandon in a white shirt and tie at a community event
01 / 02

Small details

01

Offscreen

I began dancing in ninth grade and joined Crescenta Valley’s competitive team in tenth. Our all-male squad placed fourth at nationals in 2025. I also sing with KAMA, bake, and have completed a marathon.

Dance & activities
02

Roots

My Korean name is 백규현. Growing up Korean-American and learning in a dual-language environment made both languages part of my everyday life.

That connection also shows up beyond the page: I volunteer at Korean-American heritage events through Next Generation Advocates, and help lead community music through KAMA.

Cultural volunteering
03

Learning

I started building AI projects in fourth grade and began formal research in ninth. My interests now center on machine learning, decision systems, and operations research.

I’ve taught computer science and AI since 2022. Now I’m focusing on machine learning and deep learning, and building a course around basic calculus and the core concepts.

I care about work that solves real problems, holds up to scrutiny, and is clear about its limits.

Teaching & projects
04

People

I founded HumanityOverAI to make AI ethics easier to understand. I also lead a separate five-member AI research team. Through Homemade Delights, I connect baking and mental health awareness with community service.

I’ve served as KAMA’s regional student president since 2025, organizing concerts and community performances. I’ve also stepped in as conductor for three concerts.

Community work

Contact

연락

Open to thoughtful conversations about research, collaborations, and meaningful ideas.

Let’s connect

A question worth exploring.
An idea worth building.
I’d love to hear it.

Credits

Korean name lettering: Solmoe Kim Daegeon · Summit Design ↗︎

← Projects

Published models · Ongoing development

Bori · 보리

Experimental Korean–English small language models adapted from pretrained models on free Kaggle GPUs. Bori-2 checkpoints are published; Bori-3 develops the training pipeline further.

Added Korean tokens
8,981
Bori-2 Base training steps
10,000
Instruction checkpoint · Paused
1,800

Training & checkpoints

  • Bori-2 Base adapts SmolLM2-135M with 8,981 added Korean tokens, embedding warm-up, and continued pretraining with English replay. Its published base checkpoint completed 10,000 full-pretraining steps.
  • The Bori-2 instruction checkpoint was paused at step 1,800. Its model card documents an instruction-loss masking bug and incomplete convergence.
  • The Bori-3 pipeline targets SmolLM2-360M with response-only instruction loss, a 50/50 Korean–English instruction mix, and evaluation scripts. This is development work, rather than a published benchmark result.
← Projects

Live website

Homemade Delights

Website for the student-led initiative I co-founded, supporting mental health awareness and fundraising through baking.

Co-founder & head of marketing

Survey responses
2,300+
Raised for local charities
$2,500+
Baked goods sold
500+

Beyond the website

I co-founded and co-lead a nine-person organization connecting mental health awareness with community service. I lead marketing, research-backed writing, outreach, and events.

Community work
Homemade Delights Website
← Projects

Earlier work

Bora

Small language models built and trained from scratch using free Kaggle GPU resources, with Jupyter notebooks for training.

Bora Small Language Model
← Projects

Earlier work

AI Voice Assistant

An AI Voice Assistant based on Google's Gemini. Specifically designed for macOS devices.

AI Voice Assistant
← Projects

Earlier work

AP Computer Science

Repository containing files (Labs, HW, Notes, Exercises) for a High School AP Computer Science class (1st Semester).

AP Computer Science Files
← Projects

Earlier work

California Housing Analysis

Data exploration and prediction modeling (Multiple Linear Regression) on the California Housing Dataset using Google Colab.

California Housing Analysis
← Projects

Earlier work

EmailChat

An email-based "Instant" Messaging program utilizing IMAP and SMTP to send/receive emails formatted as a chat UI.

EmailChat Program
← Projects

Earlier work

Fluent Discord Bot

A Discord translation bot using Google Translate. Features automatic translation in selected channels and user-specific embed colors.

Fluent Discord Translation Bot
← Projects

Earlier work

MNIST KNN Classifier

A nearest-neighbor classifier for handwritten MNIST digits, built in Google Colab.

MNIST KNN Classifier
← Projects

Earlier work

Random Project Pile

A collection of various smaller Python projects and scripts created in spare time (e.g., Baek Bank, Bubble Sort, File Organizer).

Random Project Pile
← Projects

Earlier work

NumPy Neural Network

A simplistic Neural Network developed from scratch using only NumPy.

Neural Network
← Projects

Earlier work

SLM RAG Project

A simple implementation of Retrieval-Augmented Generation (RAG) enhancing a Small Language Model (SmolLM-360M) with Sentence Transformer embeddings.

SLM RAG Project
← Projects

Earlier work

Student Depression Prediction

An exploratory data analysis and K-nearest-neighbor classification capstone using a student-depression dataset. A research and learning project, rather than a clinical diagnostic tool.

Student Depression Prediction
← Projects

Earlier work

YouTube Sentiment Detection

A simple Neural Network using Keras, trained on YouTube comments for text sentiment classification. Dataset provided.

YouTube Sentiment Detection
← Projects

In development

GovGuide

Making local government processes easier to understand.

← Projects

In development

Deep learning course

A machine learning and deep learning course built around basic calculus and core concepts. Not yet published.

Paper

Paper

1 / 5
Download