tensorflow中next_batch的具体使用

更新时间：2018年02月02日 09:03:03 作者：小妖精Fsky

本篇文章主要介绍了tensorflow中next_batch的具体使用，小编觉得挺不错的，现在分享给大家，也给大家做个参考。一起跟随小编过来看看吧

本文介绍了tensorflow中next_batch的具体使用，分享给大家，具体如下：

此处给出了几种不同的next_batch方法，该文章只是做出代码片段的解释，以备以后查看：

 def next_batch(self, batch_size, fake_data=False):
  """Return the next `batch_size` examples from this data set."""
  if fake_data:
   fake_image = [1] * 784
   if self.one_hot:
    fake_label = [1] + [0] * 9
   else:
    fake_label = 0
   return [fake_image for _ in xrange(batch_size)], [
     fake_label for _ in xrange(batch_size)
   ]
  start = self._index_in_epoch
  self._index_in_epoch += batch_size
  if self._index_in_epoch > self._num_examples: # epoch中的句子下标是否大于所有语料的个数，如果为True,开始新一轮的遍历
   # Finished epoch
   self._epochs_completed += 1
   # Shuffle the data
   perm = numpy.arange(self._num_examples) # arange函数用于创建等差数组
   numpy.random.shuffle(perm) # 打乱
   self._images = self._images[perm]
   self._labels = self._labels[perm]
   # Start next epoch
   start = 0
   self._index_in_epoch = batch_size
   assert batch_size <= self._num_examples
  end = self._index_in_epoch
  return self._images[start:end], self._labels[start:end]

该段代码摘自mnist.py文件，从代码第12行start = self._index_in_epoch开始解释，_index_in_epoch-1是上一次batch个图片中最后一张图片的下边，这次epoch第一张图片的下标是从 _index_in_epoch开始，最后一张图片的下标是_index_in_epoch+batch, 如果 _index_in_epoch 大于语料中图片的个数，表示这个epoch是不合适的，就算是完成了语料的一遍的遍历，所以应该对图片洗牌然后开始新一轮的语料组成batch开始

def ptb_iterator(raw_data, batch_size, num_steps):
 """Iterate on the raw PTB data.

 This generates batch_size pointers into the raw PTB data, and allows
 minibatch iteration along these pointers.

 Args:
  raw_data: one of the raw data outputs from ptb_raw_data.
  batch_size: int, the batch size.
  num_steps: int, the number of unrolls.

 Yields:
  Pairs of the batched data, each a matrix of shape [batch_size, num_steps].
  The second element of the tuple is the same data time-shifted to the
  right by one.

 Raises:
  ValueError: if batch_size or num_steps are too high.
 """
 raw_data = np.array(raw_data, dtype=np.int32)

 data_len = len(raw_data)
 batch_len = data_len // batch_size #有多少个batch
 data = np.zeros([batch_size, batch_len], dtype=np.int32) # batch_len 有多少个单词
 for i in range(batch_size): # batch_size 有多少个batch
  data[i] = raw_data[batch_len * i:batch_len * (i + 1)]

 epoch_size = (batch_len - 1) // num_steps # batch_len 是指一个batch中有多少个句子
 #epoch_size = ((len(data) // model.batch_size) - 1) // model.num_steps # // 表示整数除法
 if epoch_size == 0:
  raise ValueError("epoch_size == 0, decrease batch_size or num_steps")

 for i in range(epoch_size):
  x = data[:, i*num_steps:(i+1)*num_steps]
  y = data[:, i*num_steps+1:(i+1)*num_steps+1]
  yield (x, y)

第三种方式：

  def next(self, batch_size):
    """ Return a batch of data. When dataset end is reached, start over.
    """
    if self.batch_id == len(self.data):
      self.batch_id = 0
    batch_data = (self.data[self.batch_id:min(self.batch_id +
                         batch_size, len(self.data))])
    batch_labels = (self.labels[self.batch_id:min(self.batch_id +
                         batch_size, len(self.data))])
    batch_seqlen = (self.seqlen[self.batch_id:min(self.batch_id +
                         batch_size, len(self.data))])
    self.batch_id = min(self.batch_id + batch_size, len(self.data))
    return batch_data, batch_labels, batch_seqlen

第四种方式：

def batch_iter(sourceData, batch_size, num_epochs, shuffle=True):
  data = np.array(sourceData) # 将sourceData转换为array存储
  data_size = len(sourceData)
  num_batches_per_epoch = int(len(sourceData) / batch_size) + 1
  for epoch in range(num_epochs):
    # Shuffle the data at each epoch
    if shuffle:
      shuffle_indices = np.random.permutation(np.arange(data_size))
      shuffled_data = sourceData[shuffle_indices]
    else:
      shuffled_data = sourceData

    for batch_num in range(num_batches_per_epoch):
      start_index = batch_num * batch_size
      end_index = min((batch_num + 1) * batch_size, data_size)

      yield shuffled_data[start_index:end_index]

迭代器的用法，具体学习Python迭代器的用法

另外需要注意的是，前三种方式只是所有语料遍历一次，而最后一种方法是，所有语料遍历了num_epochs次

以上就是本文的全部内容，希望对大家的学习有所帮助，也希望大家多多支持脚本之家。

您可能感兴趣的文章:

Python3enumrate和range对比及示例详解
这篇文章主要介绍了Python3enumrate和range对比及示例详解，在Python中，enumrate和range都常用于for循环中，enumrate函数用于同时循环列表和元素，而range()函数可以生成数值范围变化的列表，而能够用于for循环即都是可迭代的,需要的朋友可以参考下
2019-07-07
python之no module named xxxx以及虚拟环境配置过程
在Python开发过程中,经常会遇到环境配置和包管理的问题,主要原因包括未安装所需包或使用虚拟环境导致的,通过pip install命令安装缺失的包是解决问题的一种方式,此外,使用虚拟环境,例如PyCharm支持的Virtualenv,可以为每个项目创建独立的运行环境
2024-10-10
Python通过4种方式实现进程数据通信
这篇文章主要介绍了Python通过4种方式实现进程数据通信,文中通过示例代码介绍的非常详细，对大家的学习或者工作具有一定的参考学习价值,需要的朋友可以参考下
2020-03-03
python2.7使用plotly绘制本地散点图和折线图
这篇文章主要为大家详细介绍了python2.7使用plotly绘制本地散点图和折线图实例，具有一定的参考价值，感兴趣的小伙伴们可以参考一下
2019-04-04
跟老齐学Python之不要红头文件(2)
在前面学习了基本的打开和建立文件之后，就可以对文件进行多种多样的操作了。请看官要注意，文件，不是什么特别的东西，就是一个对象，如同对待此前学习过的字符串、列表等一样。
2014-09-09
python 调用js的四种方式
这篇文章主要介绍了python 调用js的四种方式，帮助大家更好的理解和学习使用python，感兴趣的朋友可以了解下
2021-04-04
python脚本监控logstash进程并邮件告警实例
这篇文章主要介绍了python脚本监控logstash进程并邮件告警实例，具有很好的参考价值，希望对大家有所帮助。一起跟随小编过来看看吧
2020-04-04
python使用reportlab实现图片转换成pdf的方法
这篇文章主要介绍了python使用reportlab实现图片转换成pdf的方法,涉及Python使用reportlab模块操作图片转换的相关技巧,需要的朋友可以参考下
2015-05-05
Python7个爬虫小案例详解(附源码)上篇
这篇文章主要介绍了Python7个爬虫小案例详解（附源码）上篇，本文章内容详细，通过案例可以更好的理解爬虫的相关知识，七个例子分为了三部分，本次为上篇，共有二道题，需要的朋友可以参考下
2023-01-01
Python使用oslo.vmware管理ESXI虚拟机的示例参考
oslo.vmware是OpenStack通用框架中的一部分，主要用于实现对虚拟机的管理任务，借助oslo.vmware模块我们可以管理Vmware ESXI集群环境。
2021-06-06

tensorflow中next_batch的具体使用

相关文章

最新评论

大家感兴趣的内容

最近更新的内容

常用在线小工具