그래프를 보면 세 클래스가 꽃잎과 꽃받침의 측정값에 따라 비교적 잘 구분되는 것을 알 수 있
습니다. 이것으로 미루어보아 클래스를 잘 구분하도록 머신러닝 모델을 학습시킬 수 있을 것입
니다.
1.7.4
첫 번째 머신러닝 모델:
k
-최근접 이웃 알고리즘
이제 실제 머신러닝 모델을 만들어보겠습니다.
scikit
-
learn
은 다양한 분류 알고리즘을 제공
합니다. 여기서는 비교적 이해하기 쉬운
k
-
최근접 이웃
k
-
Nearest
Neighbors
,
k
-
NN
분류기를 사용하겠습
니다. 이 모델은 단순히 훈련 데이터를 저장하여 만들어집니다. 새로운 데이터 포인트에 대한
예측이 필요하면 알고리즘은 새 데이터 포인트에서 가장 가까운 훈련 데이터 포인트를 찾습니
다. 그런 다음 찾은 훈련 데이터의 레이블을 새 데이터 포인트의 레이블로 지정합니다.
k
-최근접 이웃 알고리즘에서
k
는 가장 가까운 이웃 ‘하나’가 아니라 훈련 데이터에서 새로운 데
이터 포인트에 가장 가까운 ‘
k
개’의 이웃을 찾는다는 뜻입니다(예를 들면 가장 가까운 세 개 혹
은 다섯 개의 이웃). 그런 다음 이 이웃들의 클래스 중 빈도가 가장 높은 클래스를 예측값으로
사용합니다. 자세한 내용은
2
장에서 살펴보기로 하고, 지금은 하나의 이웃만 사용하겠습니다.
scikit
-
learn
의 모든 머신러닝 모델은 ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month, and much more.
O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.