dython：Python資料建模寶藏庫

費弗裡發表於2021-08-15

原文網址 : https://www.cnblogs.com/feffery/p/15143481.html

儘管已經有了scikit-learn、statsmodels、seaborn等非常優秀的資料建模庫，但實際資料分析過程中常用到的一些功能場景仍然需要編寫數十行以上的程式碼才能實現。

　　而今天要給大家推薦的dython就是一款整合了諸多實用功能的資料建模工具庫，幫助我們更加高效地完成資料分析過程中的諸多工：

　　通過下面兩種方式均可完成對dython的安裝：

pip install dython

　　或：

conda install -c conda-forge dython

　　dython中目前根據功能分類劃分為以下幾個子模組：

data_utils

　　data_utils子模組整合了一些基礎性的資料探索性分析相關的API，如identify_columns_with_na()可用於快速檢查資料集中的缺失值情況：

>> df = pd.DataFrame({'col1': ['a', np.nan, 'a', 'a'], 'col2': [3, np.nan, 2, np.nan], 'col3': [1., 2., 3., 4.]})
>> identify_columns_with_na(df)
  column  na_count
1   col2         2
0   col1         1

　　identify_columns_by_type()可快速選擇資料集中具有指定資料型別的欄位：

>> df = pd.DataFrame({'col1': ['a', 'b', 'c', 'a'], 'col2': [3, 4, 2, 1], 'col3': [1., 2., 3., 4.]})
>> identify_columns_by_type(df, include=['int64', 'float64'])
['col2', 'col3']

　　one_hot_encode()可快速對陣列進行獨熱編碼：

>> one_hot_encode([1,0,5])
[[0. 1. 0. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0.]
 [0. 0. 0. 0. 0. 1.]]

　　split_hist()則可以快速繪製分組直方圖，幫助使用者快速探索資料集特徵分佈：

import pandas as pd
from sklearn import datasets
from dython.data_utils import split_hist

# Load data and convert to DataFrame
data = datasets.load_breast_cancer()
df = pd.DataFrame(data=data.data, columns=data.feature_names)
df['malignant'] = [not bool(x) for x in data.target]

# Plot histogram
split_hist(df, 'mean radius', split_by='malignant', bins=20, figsize=(15,7))

nominal

　　nominal子模組包含了一些進階的特徵相關性度量功能，例如其中的associations()可以自適應由連續型和類別型特徵混合的資料集，並自動計算出相應的Pearson、Cramer's V、Theil's U、條件熵等多樣化的係數；cluster_correlations()可以繪製出基於層次聚類的相關係數矩陣圖等實用功能：

model_utils

　　model_utils子模組包含了諸多對機器學習模型進行效能評估的工具，如ks_abc()：

from sklearn import datasets
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from dython.model_utils import ks_abc

# Load and split data
data = datasets.load_breast_cancer()
X_train, X_test, y_train, y_test = train_test_split(data.data, data.target, test_size=.5, random_state=0)

# Train model and predict
model = LogisticRegression(solver='liblinear')
model.fit(X_train, y_train)
y_pred = model.predict_proba(X_test)

# Perform KS test and compute area between curves
ks_abc(y_test, y_pred[:,1])

　　metric_graph()：

import numpy as np
from sklearn import svm, datasets
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import label_binarize
from sklearn.multiclass import OneVsRestClassifier
from dython.model_utils import metric_graph

# Load data
iris = datasets.load_iris()
X = iris.data
y = label_binarize(iris.target, classes=[0, 1, 2])

# Add noisy features
random_state = np.random.RandomState(4)
n_samples, n_features = X.shape
X = np.c_[X, random_state.randn(n_samples, 200 * n_features)]

# Train a model
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.5, random_state=0)
classifier = OneVsRestClassifier(svm.SVC(kernel='linear', probability=True, random_state=0))

# Predict
y_score = classifier.fit(X_train, y_train).predict_proba(X_test)

# Plot ROC graphs
metric_graph(y_test, y_score, 'pr', class_names=iris.target_names)

import numpy as np
from sklearn import svm, datasets
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import label_binarize
from sklearn.multiclass import OneVsRestClassifier
from dython.model_utils import metric_graph

# Load data
iris = datasets.load_iris()
X = iris.data
y = label_binarize(iris.target, classes=[0, 1, 2])

# Add noisy features
random_state = np.random.RandomState(4)
n_samples, n_features = X.shape
X = np.c_[X, random_state.randn(n_samples, 200 * n_features)]

# Train a model
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.5, random_state=0)
classifier = OneVsRestClassifier(svm.SVC(kernel='linear', probability=True, random_state=0))

# Predict
y_score = classifier.fit(X_train, y_train).predict_proba(X_test)

# Plot ROC graphs
metric_graph(y_test, y_score, 'roc', class_names=iris.target_names)

sampling

　　sampling子模組則包含了boltzmann_sampling()和weighted_sampling()兩種資料取樣方法，簡化資料建模流程。

　　dython作為一個處於快速開發迭代過程的Python庫，陸續會有更多的實用功能引入，感興趣的朋友們可以前往https://github.com/shakedzy/dython檢視更多內容或對此專案保持關注。

　　以上就是本文的全部內容，歡迎在評論區與我進行討論~

Python 資料分析入門寶藏書，選它！
2021-02-07
Python
一鍵自動化資料分析！快來看看這些寶藏工具庫
2022-07-12
雲資料建模：為資料倉儲設計資料庫
2022-06-30
資料庫
寶藏
2024-09-03
寶塔資料庫恢復 mysql資料庫丟失恢復 mysql資料庫刪除庫恢復寶塔mysql資料庫恢復
2020-11-10
資料庫MySql
Python內建模組之 re庫
2021-03-17
Python
資料建模
2024-05-08
Python數學建模-02.資料匯入
2021-05-30
Python
C#快速搭建模型資料庫SQLite操作
2019-07-22
C#模型資料庫SQLite
Python抓取淘寶IP地址資料
2019-04-26
Python
python+資料庫（三）用python對資料庫基本操作
2018-11-08
Python資料庫
python資料庫2
2021-01-03
Python資料庫
MongoDB隱藏技能：如何重新命名資料庫
2019-04-19
MongoDB資料庫
寶藏級BI資料視覺化功能|圖表聯動分析
2023-03-08
視覺化
python如何將資料插入資料庫
2021-09-11
Python資料庫
Python 操作 SQLite 資料庫
2018-12-07
PythonSQLite資料庫
Python——Reflex（資料庫使用）
2024-04-26
PythonFlex資料庫
Python操作SQLite資料庫
2019-05-14
PythonSQLite資料庫
python操作mongodb資料庫
2024-10-23
PythonMongoDB資料庫
API介面寶藏：免費好用資源分享
2023-11-03
API
Golang 學習寶庫，各種資料收集
2019-09-11
Golang
寶塔皮膚資料庫怎麼連
2021-04-02
資料庫
寶藏公司一StreamNative
2022-03-18
Python標準庫中隱藏的利器
2023-11-12
Python
推薦五款寶藏軟體，身為寶藏男孩和寶藏女孩的你，不試一下嗎？
2023-03-11
資料治理之資料梳理與建模
2024-04-25
python資料插入連線MySQL資料庫
2020-04-23
PythonMySql資料庫
BA資料建模概述 - batimes
2021-08-10
BAT
python資料庫-MySQL資料庫高階查詢操作(51)
2019-07-11
Python資料庫MySql
淘寶海量資料庫OceanBase系統架構
2022-11-24
資料庫架構
Python 資料庫騷操作 — Redis
2019-02-16
Python資料庫Redis
Python 資料庫騷操作 -- Redis
2018-11-12
Python資料庫Redis
Python 資料庫騷操作 -- MongoDB
2018-11-07
Python資料庫MongoDB
Python資料庫MongoDB騷操作
2018-11-07
Python資料庫MongoDB
Python連線SQLite資料庫
2019-01-14
PythonSQLite資料庫
Python之操作 MySQL 資料庫
2019-01-07
PythonMySql資料庫
python資料庫連線池
2020-07-07
Python資料庫
使用Python連線資料庫
2020-10-16
Python資料庫

dython：Python資料建模寶藏庫

相關文章