当前位置:网站首页>用XGBoost迭代读取数据集
用XGBoost迭代读取数据集
2022-06-27 06:35:00 【Datawhale】
Datawhale干货
来源:Coggle数据科学
在大规模数据集进行读取进行训练的过程中,迭代读取数据集是一个非常合适的选择,在Pytorch中支持迭代读取的方式。接下来我们将介绍XGBoost的迭代读取的方式。
内存数据读取
class IterLoadForDMatrix(xgb.core.DataIter):
def __init__(self, df=None, features=None, target=None, batch_size=256*1024):
self.features = features
self.target = target
self.df = df
self.batch_size = batch_size
self.batches = int( np.ceil( len(df) / self.batch_size ) )
self.it = 0 # set iterator to 0
super().__init__()
def reset(self):
'''Reset the iterator'''
self.it = 0
def next(self, input_data):
'''Yield next batch of data.'''
if self.it == self.batches:
return 0 # Return 0 when there's no more batch.
a = self.it * self.batch_size
b = min( (self.it + 1) * self.batch_size, len(self.df) )
dt = pd.DataFrame(self.df.iloc[a:b])
input_data(data=dt[self.features], label=dt[self.target]) #, weight=dt['weight'])
self.it += 1
return 1调用方法(此种方式比较适合GPU训练):
Xy_train = IterLoadForDMatrix(train.loc[train_idx], FEATURES, 'target')
dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)参考文档:
https://xgboost.readthedocs.io/en/latest/python/examples/quantile_data_iterator.html
外部数据迭代读取
class Iterator(xgboost.DataIter):
def __init__(self, svm_file_paths: List[str]):
self._file_paths = svm_file_paths
self._it = 0
super().__init__(cache_prefix=os.path.join(".", "cache"))
def next(self, input_data: Callable):
if self._it == len(self._file_paths):
# return 0 to let XGBoost know this is the end of iteration
return 0
X, y = load_svmlight_file(self._file_paths[self._it])
input_data(X, y)
self._it += 1
return 1
def reset(self):
"""Reset the iterator to its beginning"""
self._it = 0调用方法(此种方式比较适合CPU训练):
it = Iterator(["file_0.svm", "file_1.svm", "file_2.svm"])
Xy = xgboost.DMatrix(it)
# Other tree methods including ``hist`` and ``gpu_hist`` also work, but has some caveats
# as noted in following sections.
booster = xgboost.train({"tree_method": "approx"}, Xy)参考文档:
https://xgboost.readthedocs.io/en/stable/tutorials/external_memory.html

整理不易,点赞三连↓
边栏推荐
- extendible hashing
- tracepoint
- Fast realization of Bluetooth communication between MCU and mobile phone
- Interviewer: how to never migrate data and avoid hot issues by using sub database and sub table?
- Compatibility comparison between tidb and MySQL
- 0.0.0.0:x的含义
- 第 299 场周赛 第四题 6103. 从树中删除边的最小分数
- OPPO面试整理,真正的八股文,狂虐面试官
- Scala函数柯里化(Currying)
- Unsafe中的park和unpark
猜你喜欢
随机推荐
Classical cryptosystem -- substitution and replacement
TiDB 基本功能
Win10 remote connection to ECS
2018年数学建模竞赛-高温作业专用服装设计
An Empirical Evaluation of In-Memory Multi-Version Concurrency Control
Spark sql 常用时间函数
TiDB 中的数据库模式概述
NoViableAltException([email protected][2389:1: columnNameTypeOrConstraint : ( ( tableConstraint ) | ( columnNameT
0.0.0.0:x的含义
Park and unpark in unsafe
观测电机转速转矩
面试官:你天天用 Lombok,说说它什么原理?我竟然答不上来…
extendible hashing
Meaning of 0.0.0.0:x
TiDB的使用限制
Tidb database Quick Start Guide
高薪程序员&面试题精讲系列116之Redis缓存如何实现?怎么发现热key?缓存时可能存在哪些问题?
[QT dot] QT download link
HTAP 快速上手指南
Unrecognized VM option ‘‘









