중요하지 않아 보이는 특성을 제외하는 대신, 얼마나 의미 있는 특성인지를 계산해서 스케일을
조정하는 방식이 있습니다. 가장 널리 알려진 방식은
tf-idf
term
frequency
-
inverse
document
frequency
, 단어빈
도-역문서빈도
입니다.
tf
-
idf
는 말뭉치의 다른 문서보다 특정 문서에 자주 나타나는 단어에 높은 가
중치를 주는 방법입니다. 한 단어가 특정 문서에 자주 나타나고 다른 여러 문서에서는 그렇지
않다면, 그 문서의 내용을 아주 잘 설명하는 단어라고 볼 수 있습니다.
scikit
-
learn
은 두 개의
파이썬 클래스에
tf
-
idf
를 구현했습니다.
TfidfTransformer
는
CountVectorizer
가 만든 희
소 행렬을 입력받아 변환합니다.
TfidfVectorizer
는 텍스트 데이터를 입력받아
BOW
특성 추
출과
tf
-
idf
변환을 수행합니다.
17
tf
-
idf
스케일 변환 방식은 여러 변종이 있으니 위키백과를 ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month, and much more.
O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.