앞에서 설명한 이상치 프로토콜을 사용하여 이런 문제를 자동으로 경고할 수도 있습니다. 데이
터셋 관련 도메인 지식을 사용하여 데이터셋이 가능한 한 편중되지 않은 숫자 값에 대한 제한
을 적용할 수 있습니다. 예를 들어 데이터셋에 사용자의 임금이 숫자 피처로 포함된다고 가정
해보죠. 피처값의 평균이 현실적이게 강제할 수 있습니다.
편향에 관한 자세한 내용은 구글의 머신러닝 집중 과정(
https
://
oreil
.
ly
/
JtX5b
)을 참고
하세요.
4.3.5
TFDV
에서 데이터 슬라이싱하기
TFDV
를 사용하여 선택한 피처에서 데이터셋을 슬라이싱하여 데이터 편향을 확인할 수도 있
습니다. 이는
7
장에서 설명할 슬라이스 피처의 모델 성능 계산과 유사합니다. 예를 들어, 데이
터가 누락되었을 때 편향 데이터가 발생하곤 합니다. 데이터가 임의로 누락되지 않으면 데이터
셋 내의 한 사용자 그룹이 다른 사용자보다 더 자주 누락될 수 있습니다. 즉, 최종 모델을 학습
할 때 이런 그룹에서 성능이 저하됩니다.
여러 미국 주의 데이터를 예로 들어보겠습니다. 다음 코드를 사용하여 캘리포니아에서만 통계
를 얻도록 데이터를 슬라이싱할 수 있습니다.
100
살아 움직이는 머신러닝 파이프라인 설계
fromtensorflow_data_validation.utils ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month, and much more.
O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.