I have 6 large files which each of them contains a dictionary object that I saved in a hard disk using pickle function. It takes about 600 seconds to load all of them in sequential order. I want to start loading all them at the same time to speed up the process. Suppose all of them have the same size, I hope to load them in 100 seconds instead. I used multiprocessing and apply_async to load each of them separately but it runs like sequential. This is the code I used and it doesn't work. The code is for 3 of these files but it would be the same for six of them. I put the 3rd file in another hard disk to make sure the IO is not limited.


def loadMaps():    
    start = timeit.default_timer()
    procs = []
    pool = Pool(3)
    stop = timeit.default_timer()
    print('loadFiles takes in %.1f seconds' % (stop - start))

1 个解决方案



If your code is primarily limited by IO and the files are on multiple disks, you might be able to speed it up using threads:


import concurrent.futures
import pickle

def read_one(fname):
    with open(fname, 'rb') as f:
        return pickle.load(f)

def read_parallel(file_names):
    with concurrent.futures.ThreadPoolExecutor() as executor:
        futures = [executor.submit(read_one, f) for f in file_names]
        return [fut.result() for fut in futures]

The GIL will not force IO operations to run serialized because Python consistently releases it when doing IO.


Several remarks on alternatives:


  • multiprocessing is unlikely to help because, while it guarantees to do its work in multiple processes (and therefore free of the GIL), it also requires the content to be transferred between the subprocess and the main process, which takes additional time.


  • asyncio will not help you at all because it doesn't natively support asynchronous file system access (and neither do the popular OS'es). While it can emulate it with threads, the effect is the same as the code above, only with much more ceremony.


  • Neither option will speed up loading the six files by a factor of six. Consider that at least some of the time is spent creating the dictionaries, which will be serialized by the GIL. If you want to really speed up startup, a better approach is not to create the whole dictionary upfront and switch to an in-file database, possibly using the dictionary to cache access to its content.



  1. 在Python 3.x中将多个字典写入多个csv文件
  2. 如何使用python 3检查文件夹是否包含文件
  3. 如何使用未受标头影响的python导入csv文件,其中第一列为非数值
  4. python在windows中的文件路径问题
  5. 套接字。接受错误24:对许多打开的文件
  6. python如何将一个txt文件里的转化为相应字典
  7. Python之错误异常和文件处理
  8. python解析json文件读取Android permission说明
  9. python 之 logger日志 字典配置文件


  1. 如何提高android代码质量
  2. android中的颜色值
  3. 在android中玩转wcf
  4. Android入门教程(九)之-----取得手机屏幕
  5. android 监听电源键
  6. android 监听webview的超链接点击
  7. Android获取应用程序信息——PackageMana
  8. 銆婄涓€琛屼唬鐮丄ndroid銆嬬瑪璁?/h1
  9. Android(安卓)ContentResolver CallLog
  10. 《Android高级进阶》— Android 书籍