그룹을 또다시 하위 그룹으로 분할하며, 마지막에는 단 하나의 비싼 제품으로 이어지는 경로가
구축됩니다.
단일 범주형 변수를 여러 원-핫 인코딩된 열로 대체하는 방법도 있습니다. 여기서 대체된 각
열은 원래의 범주형 변수가 표현할 수 있는 모든 수준을 펼친 것입니다. 참고로 팬더스는 이를
수행하는
get
_
dummies
메서드를 제공합니다.
그러나 이 접근법이 최종 결과를 향상한다는 증거는 없습니다. 따라서 일반적으로 데이터셋
작업을 더 어렵게 만들므로 가능한 한 피해야 합니다. 이 문제는 마빈 라이트
Marvin
Wright
와 잉
케 쾨니히
Inke
König
가
2019
년에 작성한 「
Splitting
on
Categorical
Predictors
in
Random
Forests
(
https
://
oreil
.
ly
/
ojzKJ
)」 논문에서 다룬 바가 있습니다.
명목 예측 변수에 대한 표준 접근 방식은
k
개의 예측 변수 범주의 모든
2
k
-
1
-
1
2
파티션을 고려하는
것입니다. 그러나 이 지수 관계는 평가할 잠재적 분할을 많이 생성하여 계산 복잡성을 증가시키고 대
부분의 구현에서 가능한 범주 수를 제한합니다. 이진 분류 및 회귀에서는 각 분할에서 예측자 범주를
정렬하면 표준 접근 방식과 똑같은 분할이 발생합니다. 이렇게 하면 범주가
k
개인 명목 예측 변수에
대해
k
-
1
분할만 고려해야 하므로 계산 복잡성이
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month, and much more.
O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.