r/Rlanguage 커뮤니티 요약 판단 기준을 임계값 곡선으로 점검하기: 개념부터 읽기

최근 R ecosystem/community 수집 컨텍스트 `R 사용 문의 정리`에서 얻은 재현용 관심도 지표를 검열된 시간 자료 구조로 만들고, `Threshold lens` 분석 뒤 negative control과 누적분포 곡선을 사용해 결론을 다시 점검합니다.

### r/Rlanguage 커뮤니티 요약 판단 기준을 임계값 곡선으로 점검하기: 개념부터 읽기

최근 수집된 r/Rlanguage의 공개 글감 `R 사용 문의 정리`은 커뮤니티 요약에서 보이는 반복 주제을 작은 데이터 질문으로 바꿔 볼 만한 단서입니다. 이 Notebook은 원문을 복제하지 않고, 그 맥락에서 보이는 관심 방향을 재현용 예제 데이터로 바꿔 분석합니다. 수집 요약은 주제 선택의 힌트로만 사용하고, 아래 분석 값은 모두 재현 가능한 예제 데이터입니다.

참고 컨텍스트는 r/Rlanguage의 `R 사용 문의 정리` (2026-09-23)입니다. 원문 링크나 본문을 복제하지 않고, 주제 신호만 작은 분석 프레임으로 가져옵니다.

오늘은 임계값을 바꾸면 판단이 얼마나 달라지는지 살펴봅니다. 오늘의 질문은 **커뮤니티 요약에서 반복되는 관심이 실제 신호처럼 보이는가?** 입니다.

이번 글은 **검열된 시간 자료** 구조와 **개념 학습** 관점을 사용합니다. 관측 종료 전에 사건이 없었던 항목을 검열 표시와 함께 보존합니다. 처음 접하는 독자가 식과 그림을 연결할 수 있게 설명합니다.

기본 지표는 `precision = TP / (TP + FP)`와 `recall = TP / (TP + FN)`입니다. 둘 중 하나만 보면 실제 판단 비용을 놓칠 수 있습니다.

아래에서는 최근 R ecosystem/community 수집 컨텍스트 `R 사용 문의 정리`에서 얻은 재현용 관심도 지표를 만든 뒤, 브라우저 WebR에서 실행 가능한 base R 코드로 작은 분석을 진행합니다.
### 코드가 하고 있는 일

어떤 기준 이상을 표시할지 정할 때는 임계값 하나를 고정하기 전에 precision과 recall의 균형을 봅니다. 임계값을 여러 개 움직이면 민감한 구간이 드러납니다. 위 cell은 데이터 생성과 시각화를 같이 담았고, 아래 cell은 같은 객체에서 숫자 요약만 분리해 확인합니다.
### 결론을 한 번 더 흔들어 보기

주 분석 다음에는 **negative control**을 적용합니다. 관계가 없어야 하는 대조 변수를 넣어 가짜 신호 가능성을 봅니다. 결과는 **누적분포 곡선**으로 그립니다. 개별 bin 선택에 덜 민감한 ECDF로 전체 분포를 비교합니다. 또한 WebAssembly 지원 패키지 **data.table**를 실제로 불러 data.table 문법으로 segment별 요약을 계산합니다.
### 읽는 포인트

임계값 곡선은 정답 하나를 주기보다, 운영자가 감당할 오탐과 미탐의 균형점을 찾게 해줍니다. 처음 접하는 독자가 식과 그림을 연결할 수 있게 설명합니다. 오늘의 수치는 실제 운영 지표가 아니라 재현 가능한 예제 데이터이지만, 같은 분석 blueprint는 실제 로그에서도 데이터 구조와 검증 질문을 명시한 뒤 재사용할 수 있습니다.
# Threshold lens: source-context-15d4f84799b59e
set.seed(1891931503)
n <- 220
segment <- sample(c("low", "middle", "high"), n, replace = TRUE, prob = c(0.36, 0.44, 0.20))
score <- pmin(1, pmax(0, stats::rbeta(n, shape1 = 2.2 + (segment == "high"), shape2 = 2.8 - 0.4 * (segment == "high"))))
true_prob <- pmin(0.95, pmax(0.05, 0.18 + 0.62 * score + 0.10 * (segment == "high")))
outcome <- stats::rbinom(n, size = 1, prob = true_prob)
thresholds <- seq(0.15, 0.85, by = 0.05)
metrics <- data.frame(threshold = thresholds, precision = NA_real_, recall = NA_real_, flagged = NA_real_)
for (i in seq_along(thresholds)) {
  predicted <- score >= thresholds[i]
  tp <- sum(predicted & outcome == 1)
  fp <- sum(predicted & outcome == 0)
  fn <- sum(!predicted & outcome == 1)
  metrics$precision[i] <- ifelse(tp + fp == 0, NA, tp / (tp + fp))
  metrics$recall[i] <- ifelse(tp + fn == 0, NA, tp / (tp + fn))
  metrics$flagged[i] <- mean(predicted)
}
metrics$f1 <- 2 * metrics$precision * metrics$recall / (metrics$precision + metrics$recall)
best <- metrics[which.max(metrics$f1), ]

grDevices::svg("/tmp/webr_daily_threshold-lens.svg", width = 7.2, height = 4.6, bg = "white")
op <- par(mar = c(4.5, 4.7, 3.2, 4.2), bg = "white")
plot(metrics$threshold, metrics$precision, type = "b", pch = 19, col = "#0f766e", ylim = c(0, 1), xlab = "threshold", ylab = "precision / recall", main = "R ecosystem source pulse 15d4")
lines(metrics$threshold, metrics$recall, type = "b", pch = 21, bg = "#f97316", col = "#f97316")
lines(metrics$threshold, metrics$f1, type = "b", pch = 17, col = "#111827")
abline(v = best$threshold, col = "#111827", lwd = 2, lty = 2)
legend("bottomleft", legend = c("precision", "recall", "F1", "best threshold"), col = c("#0f766e", "#f97316", "#111827", "#111827"), pch = c(19, 21, 17, NA), lty = c(1, 1, 1, 2), bty = "n")
par(op)
grDevices::dev.off()
# Threshold lens summary
cat("r/Rlanguage 커뮤니티 요약", "threshold lens\n")
cat("best threshold:", round(best$threshold, 2), "\n")
cat("precision at best:", round(best$precision, 3), "\n")
cat("recall at best:", round(best$recall, 3), "\n")
cat("flagged share at best:", paste0(round(100 * best$flagged, 1), "%"), "\n")
# Blueprint validation: 744e671f5ba4e09232105463af07e58d9d4c61e7309234f1c352673810ca87d9
set.seed(1891939422)
time_index <- seq_len(180)
segment <- rep(c("A", "B", "C"), length.out = length(time_index))
event_time <- rexp(length(time_index), rate = 0.08 + 0.02 * (segment == "C"))
censor_time <- runif(length(time_index), 4, 22)
observed <- event_time <= censor_time
probe <- pmin(event_time, censor_time)
lens_values <- replicate(320, suppressWarnings(cor(probe, sample(time_index))))
lens_axis <- seq_along(lens_values)
lens_summary <- mean(abs(lens_values) > 0.2, na.rm = TRUE)
lens_values <- as.numeric(lens_values)
lens_values <- lens_values[is.finite(lens_values)]
if (!length(lens_values)) lens_values <- 0
topic_color <- "#0f766e"
topic_accent <- "#f97316"
visual_title <- "negative control · 누적분포 곡선"
grDevices::svg("/tmp/webr_daily_validation_744e671f5ba4.svg", width = 7.2, height = 4.6, bg = 'white')
op <- par(mar = c(4.6, 4.8, 3.2, 1.2), bg = 'white')
plot(stats::ecdf(lens_values), verticals = TRUE, do.points = FALSE, col = topic_color, lwd = 3, main = visual_title, xlab = 'validation value', ylab = 'cumulative probability')
par(op)
grDevices::dev.off()
pkg_object <- data.table::data.table(probe = probe, segment = segment)[, .(center = median(probe)), by = segment]; package_result <- mean(pkg_object$center)
cat('WebAssembly package:', "data.table", as.character(utils::packageVersion("data.table")), '\n')
cat('package calculation:', round(as.numeric(package_result)[1], 4), '\n')
cat('data design:', "검열된 시간 자료", '\n')
cat('validation lens:', "negative control", '\n')
cat('visual grammar:', "누적분포 곡선", '\n')
cat('validation summary:', round(lens_summary, 4), '\n')
r/Rlanguage 커뮤니티 요약 threshold lens
best threshold: 0.15 
precision at best: 0.562 
recall at best: 0.975 
flagged share at best: 95.5%