写在前面

此系列文章是我自己在课余结合课堂内容对入门机器学习实际操作的历程,从这篇文章开始进入基础部分的总结与分享。我个人认为机器学习基础分为三部分:python基础(更好的看懂现成的代码,有能力上手实践一番)、机器学习概念基础(便于理解一些博文中的专业术语,针对性的理解代码中某一个点的作用)、数学基础(线性代数:如矩阵运算;运筹于优化:如梯度下降法;概率论与数理统计:如贝叶斯定理),当然还涉及到各种学习数据来源(如爬虫、sql)。对于这个系列会涉及到哪些内容感兴趣的同学,也欢迎大家访问《从0开始机器学习--0.序》~~(❤ ω ❤)

这篇文章的主题是python基础,由于python是一个发展已经较为全面的高级语言了,加之各种库更加丰满了其范围,内容无法在一篇博文中详尽的囊括。鉴于此系列主题为《入门机器学习》,故在这篇文章中以分享机器学习密切相关的内容为主。

  • 1.python基础;
  • 2.ai模型概念+基础;
  • 3.数据预处理;
  • 4.机器学习模型--1.编码(嵌入);2.聚类;3.降维;4.回归(预测);5.分类;
  • 5.正则化技术;
  • 6.神经网络模型--1.概念+基础;2.几种常见的神经网络模型;
  • 7.对回归、分类模型的评价方式;
  • 8.简单强化学习概念;
  • 9.几种常见的启发式算法及应用场景;
  • 10.机器学习延申应用-数据分析相关内容--1.A/B Test;2.辛普森悖论;3.蒙特卡洛模拟;
  • 11.数据挖掘--关联规则挖掘
  • 12.数学建模--决策分析方法,评价模型
  • 13.主动学习(半监督学习)
  • 以及其他的与人工智能相关的学习经历,如数据挖掘、计算机视觉-OCR光学字符识别、大模型等。

合集链接https://blog.csdn.net/m0_73752612/category_12799244.html?fromshare=blogcolumn&sharetype=blogcolumn&sharerId=12799244&sharerefer=PC&sharesource=m0_73752612&sharefrom=from_link


目录

写在前面

python简介

python是什么

基础语法

面经:python中的基础数据结构

面经:for和while循环的区别

1. 语法 & 使用场景

2. 循环结束的控制

python中常用于数据处理的库和指标

numpy

pandas

面经:Series 和 DataFrame

面经:python中pandas库类似sql的操作

如何拼接dataframe

如何计算平均值和中位数

实现 SQL 中的 WHERE 操作

如何删除某一列

实现 SQL 中的 GROUP BY 操作

re

matplotlib

python代码中常出现的一些其他语法现象

类型提示(type hints)

“@” -- 语法糖

lambda函数

“Polars” -- python库

多进程

multiprocessing 模块

ThreadPoolExecutor

ProcessPoolExecutor

异步处理

gc.collect() -- 释放内存的函数

总结


python简介

python是什么

Python是一种高级编程语言,具有简洁易读的语法,适合快速开发。它支持多种编程范式,包括面向对象、函数式编程等。Python有广泛的标准库,适合处理数据分析、机器学习、Web开发等各种应用。由于其灵活性和易用性,Python在科研、金融、人工智能等领域非常流行。

对于如何安装python、配置环境、亦或是安装anaconda或jupyter notebook、选择适合自己的编译工具和解释器都有很多现成的博文教学,就不在这里赘述了,当然有需要也可以评论或者私信我。

这里我想特别提醒一点的是,python所有的库文件,就比如下文会提到的numpy、pandas,默认的安装路径都在系统盘C盘中,随着不断的安装各种库,系统盘的空间很有可能会不堪重负。大家可以找一下“C:\Users\<你的用户名>\AppData\Local\Programs\Python\Python<版本号>\Lib\site-packages”和“C:\Users\<你的用户名>\AppData\Local\Packages\Python\Python<版本号>\Lib\site-packages”。几种主流的解决方法分别是修改pip的安装路径、更改环境变量中site-packages的位置、使用虚拟环境venv。我个人目前是采取了更改环境变量中的位置以释放之前占用的C盘空间,同时随着实验和项目的复杂程度不断升级,使用venv以安装规定版本的库(也可以anaconda一步搞定)。这一点在很多python安装的教程中都不会提及。

基础语法

# -*- coding: utf-8 -*-

import numpy as np
##########################################################

# 复制和简单的切片操作

print("hello world!",
      100,300)
print(3>2)        #print bool
a=np.arange(4)
b=a
c=a.copy()        #deep copy (只会复制)
print('b is a ',b is a)
print('a is b',a is b)
print('a is c',a is c)
a[0]=11
print('b',b)      #b will change with a 
print('c',c)      #c will not change with a 
b[1:4]=[22,33,44]
print(a)          #a will also change with b
print( )
##########################################################

# 演示定义全局变量和函数作用域

a=None
def fun():
    global a
    a="haha"  #change the 'x' inside of a function
fun()         #use function in order to change the value of a
print(a)
print( )
##########################################################

# 向文件中写入或读出文本

text='This is try\nThis is another try/n'
file=open("myfile.txt",'w')
file.write(text)
file.close()
file=open("myfile.txt",'r')
#content1=file.readline()
#print(content1)
content2=file.readlines()
print(content2)
file.close()
# file.append()
# content=file.read()
# content=file.readlines()
print( )

# 读取数据表格

import pandas as pd
# 读取Excel文件
file_path = '文件路径.xlsx'  # 替换为你自己的文件路径
df = pd.read_excel(file_path)
# 查看前5行数据
print(df.head())

# 使用with open读取二进制Excel文件
file_path = '文件路径.xlsx'  # 替换为你自己的文件路径
with open(file_path, 'rb') as f:
    # 读取Excel内容
    df = pd.read_excel(f)
# 查看前5行数据
print(df.head())
##########################################################
augmented_sentence='shushu快放暑假, 学生们等不及le'
print(augmented_sentence)
augmented_sentence_sent = list(augmented_sentence)
print(augmented_sentence_sent)
sent=['aaa','bbb','ccc']
a=(sent,int(3))+(augmented_sentence_sent,int(5))
print(a)
dataset = []
b=dataset.append((augmented_sentence_sent,1))
print(b) #only append to dataset itself, b=NONE
b=dataset.append((sent,2))
print(dataset)
print( )
##########################################################

# python的类

class Calculator():      #in()is the father class父类
    name="计算器"        #inherent attribute
    price=10000
    def __init__(self,weight,height,price): #input attribute; like constructor in c++; all x that have a default must be put at behind x that dont have a defult
        self.w=weight
        self.h=height
        self.p=price
        self.add(1,2)      #be an attributor, will output result once run the program
    def add(self,x,y):     #self(->*this)  must have self  different from def in outside
        print(self.name)
        result=x+y
        print(result)
calcul=Calculator(weight=20,height=10,price=200)
print(calcul.p)          #print the data given by init(200
print(calcul.price)      #print the data given by class(10000
calcul.add(11,11) 
print( )
##########################################################
a_list=[1,3,5,2,4,6,10,200]
a_list.append(999)
print(a_list)
a_list.remove(3)
print(a_list)
a_list.sort()
print(a_list)
print(a_list[3:5])
print(a_list[-3:-1]) #also[,)
print( )
###########################################################

# 其他基础语法

#input() return a string,必须做类型转换才能输入int
age = int(input('please input a number:'))  
print('This is age:',age)
if age >= 18:
    print('adult')
else:
    print('teenager')

面经:python中的基础数据结构

1. 列表(List):

  • 动态数组,支持索引、切片操作,可以存储不同类型的数据。
  • 例如:lst = [1, 2, 'apple', [3, 4]]
  • 常用操作:添加(append),删除(remove),排序(sort)等。

2. 元组(Tuple):

  • 不可变的有序集合,一旦定义无法修改元素。不要求元素具有唯一性。
  • 例如:tpl = (1, 2, 'apple')
  • 常用于需要不可变性的数据结构,效率比列表稍高。

3. 字典(Dictionary):

  • 键值对的无序集合,键必须是唯一且不可变的,值可以是任意类型。
  • 例如:dct = {'name': 'Alice', 'age': 25}
  • 常用操作:访问(dct['name']),添加或修改(dct['key'] = value),删除(del dct['key'])。

4. 集合(Set):

  • 无序且元素唯一的集合。(v.s元组:元组不要求元素唯一)
  • 例如:st = {1, 2, 3, 4}
  • 常用于去重和集合操作(交集、并集、差集)。
  • 操作:add、remove、union、intersection 等。

5. 字符串(String):

  • 不可变的字符序列,可以看作是字符的数组,支持索引、切片等操作。
  • 例如:s = "hello"
  • 常用操作:strip、replace、split、join 等。

6. 队列(Queue):

  • 通过 collections.deque 实现的双向队列,可以高效地在两端添加或删除元素。
  • 例如:from collections import deque; q = deque([1, 2, 3])
  • 常用操作:append、appendleft、pop、popleft。

7. 堆(Heap):

  • 通过 heapq 模块实现的最小堆数据结构,用于优先队列。
  • 例如:import heapq; heap = [1, 3, 5]; heapq.heapify(heap)

8. 栈(Stack):

  • Python 没有专门的栈数据结构,但可以使用列表或 deque 模拟栈。
  • 列表模拟栈:lst.append(1),lst.pop()。

面经:for和while循环的区别

1. 语法 & 使用场景

   for 循环:

  • 适用于已知次数或可以迭代的数据结构(如列表、元组、字符串、字典等)。
  • 每次循环都会从可迭代对象中自动取出下一个元素,直到迭代完成。
  • 典型的用法是遍历集合或范围。
for element in iterable:
    # 执行操作

for i in range(5):  # 遍历 0 到 4
    # 执行操作

   while 循环:

  • 适用于不确定循环次数的场景,通常依赖于某个条件是否为真来决定是否继续循环。
  • 只要条件为 True,循环就会继续执行,直到条件变为 False。
while condition:
    # 执行操作

count = 0
while count < 5:  # 当 count 小于 5 时,继续循环
    # 执行操作
    count += 1

2. 循环结束的控制

   for 循环:

  • 在遍历完可迭代对象或执行预定次数后,自动结束。
  • 不需要额外设置条件来停止循环。

   while 循环:

  • 必须明确地设置一个终止条件,否则容易进入无限循环。

python中常用于数据处理的库和指标

numpy

数组操作 切片、累加、累减、最大最小平均、维度......

import numpy as np

array=np.array([[1,4,6],[2,5,3]],dtype=int) #create an numpy array
print('array',array)
print('array[1]',array[1])
print('array[0][2] or array[0,2]',array[0][2])
print('array[0,:]',array[0,:])
print('array[0,1:3]',array[0,1:3])
print('np.cumsum(a)',np.cumsum(array)) # add by add, return a one dimension array
print('np.diff(a)',np.diff(array))     #minus between every one, return a matrix(n-1,n-1)
print('np.sum(a)',np.sum(array))       #return a num
print('np.min(a)',np.min(array))
print('np.argmin(a)',np.argmin(array)) # the index of min
print('np.min(a,axis=0)',np.min(array,axis=0)) #find min in every column, result as a row
print('np.max(a,axis=1)',np.max(array,axis=1)) #find max in every row, result as a column
print('a.mean() or np.mean(a) or np.average(a)',np.mean(array))  #average
print('np.median(a)',np.median(array)) #the data in the middle position in one dimension
print('dtype',array.dtype)
print('number of dim',array.ndim)
print('shape',array.shape)
print('size(num of element)',array.size)
print(np.nonzero(array)) #print index of data no zero
print('np.transpose(array) or array.T',np.transpose(array)) #T 转置, can only be used in matrix, not in 序列(sequence) 
for row in array: #print row
    print(row)
for column in array.T:   #print column in transpose
    print(column)
print('np.clip(array,3,5)',np.clip(array,3,5)) #make less than 3 =3; bigger than 5=5
print('np.sort(array)',np.sort(array))         #sort in each dimension
a=np.zeros((3,4)) #create a array 3*4 with all 0
print('a',a) 
b=np.ones((4,2))*3 #the data num is 1*3(=3)
print('b',b)
c=np.empty((3,4)) #create a array randomly(no initialize)
print('c',c)
d=np.arange(10,20,2) #create a array,[10,20),step=2
print('d',d)
e=np.arange(12).reshape((3,4)) #reshape's arguments must be a tuple,so needs"(())"
print('e',e)
print('e.flat',e.flat) #e.flat returns a 迭代器
print('e.flatten',e.flatten())
print('np.split(e,2,axis=1)',np.split(e,2,axis=1)) #devide rows into 2 parts
print('np.split(e,3,axis=0)',np.split(e,3,axis=0)) #parts num only can be divided all by parts
print('np.array_split(e,3,axis=1)',np.array_split(e,3,axis=1)) #split into parts with no strict
print('np.vsplit(e,3)',np.vsplit(e,3))
print('np.hsplit(e,2)',np.hsplit(e,2)) #np.split(e,2,axis=1) #devide rows into 2 parts
for item in e.flat:
    print(item)
f=np.linspace(0,24,12).reshape((3,4)) #divide [0,24]into 11(+1) parts and reshape into3*4
print('f',f)
g=e-f
print('g',g)
h=d**3 #pow(d,3)
print('h',h)
i=np.dot(e,b)   #matrix multiplication
print('i=np.dot(e,b)',i)
i=e.dot(b)
print('i=e.dot(b)',i)
i=e*f           #only mult corresponse position
print('i=e*f',i)
j=10*np.sin(i)  #rad angle
print('j',j)
print('j>0',j>0)
k=np.random.random((2,4)) #in module random use function random,shape=(2,4),data=[0,1]
print('k',k)
l1=np.array([1,11,111,1111])
l2=np.array([2,22,222,2222])
print('l1',l1)
print('l2',l2)
print('l2.shape',l2.shape)
l1l2=np.vstack((l1,l2)) #add rows in different column
l2l1=np.hstack((l1,l2)) #add columns in one row
l1l1l2l1=np.vstack((l1,l1,l2,l1)) #add several matrixs in columns in order
l1l1l2l1_2=np.concatenate((l1,l1,l2,l1),axis=0) #axis = 0  add several matrixs in the same row dimensions
print('np.vstack((l1,l2))',l1l2)
print('np.vstack((l1,l2)).shape',l1l2.shape)
print('np.hstack((l1,l2))',l2l1)
print('np.hstack((l1,l2)).shape',l2l1.shape)
print('np.vstack((l1,l1,l2,l1))',l1l1l2l1)
print('np.hstack((l1,l1,l2,l1))',np.hstack((l1,l1,l2,l1)))
print('np.concatenate((l1,l1,l2,l1),axis=0)',l1l1l2l1_2)
l3=np.array([1,11,111,1111])[:,np.newaxis] #because no T in sequence, use this method to make it into a columns' matrix
l4=np.array([1,11,111,1111])[np.newaxis,:] #use this method to make it into a rows' matrix instead of a sequence,(duoyige[])
l3l3l3=np.concatenate((l3,l3,l3),axis=1) #axis = 1  add several matrixs in different column dimensions
print('l1[:,np.newaxis]',l3)
print('l1[np.newaxis,:]',l4)
print('np.concatenate((l3,l3,l3),axis=1)',l3l3l3) 
x=np.eye(4) # 4*4单位矩阵
print(x)
print( )

pandas

dataframe、定位元素、dropna/drop_duplication、描述、转置(T)、拼接(concat/join)......

import pandas as pd
s=pd.Series([1,3,6,np.nan,44,1]) #create a thing like a dictionary, np.nan=NONE/NaN, pd.series is create one column with auto index 0.1.2......
print('pd.Series([1,3,6,np.nan,44,1])   :   ')
print(s)
dates=pd.date_range('20240403',periods=6)
print('pd.date_range(\'20240403\',periods=6)   :   ')
print(dates)
df1=pd.DataFrame(np.arange(12).reshape((3,4)))
print(df1)
df=pd.DataFrame(np.random.randn(6,4),index=dates,columns=['a','b','c','d'])#dataframe similiar to a 2 dimension numpy; data is from np.random.randn, which is 正态分布;shape(6,4);line index=index;row index=column
print('df=pd.DataFrame(np.random.randn(6,4),index=dates,columns=[\'a\',\'b\',\'c\',\'d\'])   :   ')
print(df)
print('df[df.a>1]   :') #boolean index
print(df[df.a>1])       # compare 1 with the data in row "a", and display the whole dataframe rows(bcd...)
print('df.a[df.a>1]   :')
print(df.a[df.a>1])     #display the 'a' dataframe meet the requirement(no bcd...)
print('df.a or df[\'A\']   :   ')
print(df.a)             #or print(df['A'])
print('df[0:3]   :   ')
print(df[0:3]) #lines 0.1.2
print('df.loc[\'20240404\']   :')
print(df.loc['20240404'])#select by label ".loc"  only by str(label)
print('df.loc[:,[\'a\',\'b\']]   :')
print(df.loc[:,['a','b']])
print('df.iloc[3:5,1:3]   :')    # lines 0.1.2.(3.4) columns 0.(1.2)
print(df.iloc[3:5,1:3])          #select by position ".loc"   only by num(position)
print('df.iloc[[1,2,5],1:3]   :')#select lines one by one,0.(1).(2).3.4.(5)
print(df.iloc[[1,2,5],1:3])
df.select_dtypes(include=[np.number])  # 筛选数值型列,include=[np.number]:指定筛选所有数值类型的列(包括 int、float 等)。返回一个新的数据框 numeric_features,仅包含数值型列。
df.iloc[2,2]=5211
df.loc['20240403','a']=1331
df.b[df.a>0]=14
df['add1']=np.array([[1],[111],[1],[111],[1],[111]]) #add some new columns with labels to dataframe
df['add2']=np.array([6,66,666,666,66,6])
df['add3']=np.array('add')
df['add4']=pd.Series([1,3,6,np.nan,44,1],index=df.index) #while adding a pd.series, must have index=df.index, in order to add data in required lines
print('df.iloc[2,2]=5211 && df.loc[\'20240403\',\'a\']=1331 && df.b[df.a>0]=14') #modify data in specify position 0.1.(2)
print(df)
print('np.any(df.isna()==True)   :')
print(np.any(df.isna()==True))
print('df.isnull()   :')           #return true or false in every position
print(df.isnull())
print('df.dropna(axis=0,how=\'any\')')
print(df.dropna(axis=0,how='any')) #axis=0: drop lines; how={"any", "all"}(find NaN)
df.fillna(value=0) #change every NaN into 'value'(0 here)
df.fillna(列.mode()[0]) # 众数填充,[0]表示取当前列第一多的众数的值
df2=pd.DataFrame({'A':[1,11,111,1111], 
                  'B':pd.Timestamp('20240520'), ##if(240520),then it will print 2020-05-24  #one data structure only be able to show time size(month1-12......)
                  'C':pd.Series(1,index=list(range(4)),dtype='float64'), #create a pd.Series dataset,contains 4 "1"s, and index from 0 to 3
                  'D':np.array([3]*4,dtype='int32'),
                  'E':pd.Categorical(["test",'train','test','train']), 
                  'F':'foo'},
                 index=['k0','k1','k2','k3'])#will copy n lines in order to fit the size itself if only given one dimension for 'A '
print(df2)
print(df2.dtypes)#dtype same as numpy(dimension types)
print('df2.index, df2.columns  :')
print(df2.index, df2.columns)#index 行标; columns 列标
print('df2.describe()  :')
print(df2.describe())        #show some characters of the number type data ##std:方差
print('df2.T  :')
print(df2.T)                 #pd.DataFrame can also be transpose
print('df2.sort_index(axis=1,ascending=False)  :')
print(df2.sort_index(axis=1,ascending=False))#operating on the columns'index ##ascending:递增,上升
print('df2.sort_values(by=\'E\')  :')
print(df2.sort_values(by='E'))
a = pd.Categorical(['a','b','c','a','b','c'], ordered=True, categories=['c', 'a'])#categories:the rule of sort;ordered:is sorted or not
print('pd.Categorical([\'a\',\'b\',\'c\',\'a\',\'b\',\'c\'], ordered=True, categories=[\'c\', \'b\', \'a\'])  :')
print(a)
b = pd.Categorical(['a','b','c','a','b','c'])
print(b)
print( )
# datacsv=pd.read_csv('path to the .csv') #it will add a line index (0.1.2.3...)default and return as a dataframe table
# datacsv.to_pickle('filename')# there are many kinds of read_xxx and to_xxx
df00=pd.DataFrame(np.ones((3,4))*0,columns=['a','b','c','d'])
df11=pd.DataFrame(np.ones((3,4))*1,columns=['a','b','c','d'])
df22=pd.DataFrame(np.ones((3,4))*2,columns=['a','b','c','d'])
res=pd.concat([df00,df11,df22],axis=0)          #vertical, operate on lines
print('pd.concat([df00,df11,df22],axis=0)   :') #index of lines 0.1.2.0.1.2.....
print(res)
print('pd.concat([df00,df11,df22],axis=0,ignore_index=True)   :')
resii=pd.concat([df00,df11,df22],axis=0,ignore_index=True) #make index of lines 0.1.2.3.4.5......
print(resii)
df001=pd.DataFrame(np.ones((3,4))*0,columns=['a','b','c','d'],index=[1,2,3])
df111=pd.DataFrame(np.ones((3,4))*1,columns=['b','c','d','e'],index=[2,3,4])
res21=pd.concat([df001,df111],join='inner') #find column's label's 交集
print('pd.concat([df001,df111],join=\'inner\')   :')
print(res21)
res22=pd.concat([df001,df111],join='outer') #find column's label's 并集,and fix with NaN
print('pd.concat([df001,df111],join=\'outer\')   :')
print(res22)
# restry=pd.merge(df001,df111,on=['a','b'],suffixes=['_a','_b'],how='inner',indicator=True)     # "on=[]",key of hebing; suffixes:to tell different columns' names in the hebing table, if before hebing data has two columns with the same name ; how=['outer','inner','right','left']; indicator:show where the data is from
print( )

面经:Series 和 DataFrame

Series 和 DataFrame 都是 Pandas 的核心数据结构,它们不一样,也不是普通的列表。

  • Series ≈ 带标签的列(像 Excel 中的一列),类似一维数组,但有索引。
  • DataFrame ≈ 多个 Series 的集合(像整个 Excel 表格)

面经:python中pandas库类似sql的操作

如何拼接dataframe

在 pandas 中,可以使用多种方式连接两个 DataFrame,类似 SQL 的 JOIN 操作。常见的方法包括 merge()、join() 和 concat()。

1. merge() 方法(类似 SQL 中的 JOIN,指定left即可为left join)

  • merge() 是 pandas 中最常用的用于连接两个 DataFrame 的方法,类似于 SQL 的 JOIN 操作。可以指定连接的方式,如 inner(内连接)、outer(外连接)、left(左连接)和 right(右连接)。
pd.merge(df1, df2, how='join_type', on='key_column')
  • 该方法默认为内连接,只保留在两个 DataFrame 中都有的键 key。

2. join() 方法

  • join() 是基于索引的连接操作。它可以连接两个 DataFrame,但是使用的是索引,而不是列作为键。
df1.join(df2, how='join_type', lsuffix='_left', rsuffix='_right')
  • how: 可以指定连接类型,如 left(左连接,默认值)、right(右连接)、inner 或 outer
  • lsuffix 和 rsuffix(可选):如果左右 DataFrame 中有相同的列名,可以通过这些参数为列名添加后缀

3. concat() 方法

  • concat() 用于沿着某一轴(行或列)连接多个 DataFrame,可以选择是否按列或按行进行连接。
pd.concat([df1, df2], axis=0 or 1, join='join_type')
  • axis: 指定连接的方向,0 表示按行连接(上下合并),1 表示按列连接(左右合并)
  • join: 指定连接方式,默认是 outer,可以设置为 inner
如何计算平均值和中位数
df['column_name'].mean()
df['column_name'].median()
实现 SQL 中的 WHERE 操作

在 pandas 中,WHERE 可以通过条件筛选来实现。可以通过 df[df['column_name'] <condition>] 来进行筛选。

filtered_df = df[df['A'] > 3]  # 筛选出列 A 中大于 3 的行
如何删除某一列
df.drop('column_name', axis=1, inplace=True)
  • axis=1 表示按列删除
  • inplace=True 表示原地修改 DataFrame,即删除列后不用重新赋值
实现 SQL 中的 GROUP BY 操作

pandas 中的 groupby() 方法可以实现 SQL 中的 GROUP BY 功能。常常与聚合函数(如 sum、mean、count)一起使用。

# 按 A 列的值进行分组,并对 B 列求和
data = {'A': ['group1', 'group2', 'group1', 'group2'], 'B': [10, 20, 30, 40]}
df = pd.DataFrame(data)
# 按 A 列分组,并对 B 列求和
grouped_df = df.groupby('A')['B'].sum()
print(grouped_df)

以上代码输出: 

A
group1    40
group2    60
Name: B, dtype: int64
  • .agg() 可以一次性应用多个聚合函数,并且可以针对不同的列指定不同的函数。它既可以传入单个函数,也可以传入多个函数(以列表形式)。
  • .agg() 不仅可以使用内置的聚合函数,还可以传入自定义函数。
df.groupby('group_column').agg({
    'column1': 'sum',
    'column2': ['mean', 'max'],
    'column3': lambda x: x.max() - x.min(),  # 自定义函数:最大值减去最小值
    'column4': 'std'  # 平均值和标准差
})

re

字符串匹配工具

pattern1="cat"
pattern2="bird"
string="dog runs to cat"
print(pattern1 in string)
print(pattern2 in string)

import re
print(re.search(pattern1,string))
print(re.search(pattern2,string))

#mutiple patterns(run or ran)
ptn=r"r[au]n" #在最前面加r是使它变成一个正则表达式
print(re.search(ptn,string))
print(re.search(r"r[0-9]n","dogs r2n to a cat"))
print(re.search(r"r[0-9a-zA-Z]n","dogs run r2n to a cat")) #可以多个模糊匹配,但只会返回第一个匹配到的结果
#\d:数字 \D:数字的反面即不是数字 \s:所有的空格(\t\n\r\f\v) \S:所有不是空格的东西 \w:所有字母数字和“_”[a-zA-Z0-9]
print(re.search(r"r\dn","run r2n"))
print(re.search(r"r\Dn","run r2n"))
#\b:空格(只能匹配贴着文字的空格) \B:也是匹配空格,但不是贴着文字的空格
print(re.search(r"\brun\b"," run "))
print(re.search(r"\b run \b","  run  "))#会返回None
print(re.search(r"\B run \B","  run  "))
print(re.search(r"\Brun\B"," run "))#会返回None
#\\:匹配“\”  .:匹配任何,除了/n  ^:出现在句首才能匹配到  $:出现在句尾才能匹配到  ?:may or may not
print(re.search(r"Mon(day)?","Is today Monday?"))
print(re.search(r"Mon(day)?","Is today Mon?"))
print(re.search(r"r.n"," r{n "))
print(re.search(r"^love","love you"))
print(re.search(r"^love","i love"))
#flags=re.M:匹配在多行中
string="""
a is apple.
b is boy
"""
print(re.search(r"^b",string))#会返回none 因为第一句的句首不是b
print(re.search(r"^b",string,flags=re.M))#意思是在多行中查找
#*:匹配n次 n=[0,∞)  +:匹配n次 n=[1,∞)  {n,m}:可选次数[n,m]次
print(re.search(r"ab*","a"))
print(re.search(r"ab*","abbbbbb b"))
print(re.search(r"ab+","a"))#会返回None 因为0次b
print(re.search(r"ab+","abbbbbb b"))
print(re.search(r"ab{2,5}","abbbbb b"))
#group
match=re.search(r"(\d+), Data:(.+)", "ID:22122842, Data:04/05/08/ONE")#(\d+)表示至少一个数字,(.+)表示至少一个字符  ()就是一个组
print(match.group())
print(match.group(1)) #在group()里给数字,就是返回匹配的第n个括号里的东西
print(match.group(2))
print("next match")
match=re.search(r"(?P<id>\d+), Data:(?P<data>.+)", "ID:22122842, Data:04/05/08/ONE") #?P<name> 给组加名字name,通过group(名字)查找,而不用数第几个了
print(match.group('id'))
print(match.group(1)) #和上一行输出同一个东西
print(match.group('data'))
print( )
#findall  #|:or
print(re.findall(r"r[ua]n\w*","run ran runs ren")) #\w* 返回匹配r[ua]n后有[0,∞)个数字/字母
print(re.findall(r"r(u|a)n","run ran runs ren")) #只会匹配括号中的东西,即这里的u或a
#sub:replace   split
print( )
print(re.sub(r"do(es)?","did","I do miss you, so does he")) #re.sub()默认情况下会替换所有匹配的子字符串,即已经与findall结合了
print("findall和sub结合")
text = "I do miss you, so does he"
pattern = r"do(es)?"
replacement = "did"
matches = re.findall(pattern, text)
print("matches = re.findall(pattern, text)")
print(matches)
replaced_text = text
for match in matches:
    replaced_text = re.sub(pattern, replacement, replaced_text, count=1)
print(replaced_text)
print(re.split(r"[,;.]","a,b;c;.d e")) #一定要加[],代表去掉[]中的内容做分割 该例子中的结果会有一个'',是;.中间分割出来的内容
#compile
compiled_re=re.compile(r"r[ua]n")
print(compiled_re.search('dog can run fast')) #相当于把re.search(r"a","b")中的r"a"替换为提前编译完的compiled_re
print( )

matplotlib

这是一个用于绘图的库,可以可视化数据,如柱状图、散点图、折线图,常用于与回归模型一起使用。这里只列出几个较为重要的操作。

import matplotlib.pyplot as plt 
plt.rcParams['font.sans-serif'] = ['SimHei']  # 使用黑体,解决中文显示问题
plt.rcParams['axes.unicode_minus'] = False    # 解决负号显示问题
pd.set_option('display.max_columns',None)     # 显示全部列
x=np.linspace(-10,10,50)
y1=x
y2=x**2
plt.figure()   #create a window
plt.plot(x,y1) #define 横坐标 and 纵坐标
plt.figure(num=3,figsize=(12,5))
plt.plot(x,y2)
plt.plot(x,y1,color="red",linewidth=1.0,linestyle='--') # many lines can be shown in one figure
plt.xlim((-1,2))
plt.ylim((-2,2))     #define the range of x and y in plt
plt.xlabel('I am x')
plt.ylabel('I am y') #define the label of x and y
new_ticks=np.linspace(-1,2,5)
plt.xticks(new_ticks)#define new 坐标 of x 轴
plt.yticks([-2,-0.5,0,1.58,2],[r'$really\ bad\ \alpha$','bad','normal','good','really good'])              #define new 坐标 of y with labels 自定义(match one by one)
# gca = 'get current axis'
ax = plt.gca()
ax.spines['right'].set_color('none') #spines four 边框,set none,in order to move the position of x and y 轴
ax.spines['top'].set_color('none')
ax.xaxis.set_ticks_position('bottom') # use bottom spine to use as x 轴
# ACCEPTS: [ 'top' | 'bottom' | 'both' | 'default' | 'none' ]
ax.spines['bottom'].set_position(('data', 0)) #depend on data and asign data=0 is the yuandian of the plt(move x zhou to y zhou's '0')
# the 1st is in 'outward' | 'axes' | 'data'
# axes: percentage of y axis (move x to this persentage of y)
# data: depend on y data
ax.yaxis.set_ticks_position('left')
# ACCEPTS: [ 'left' | 'right' | 'both' | 'default' | 'none' ]
ax.spines['left'].set_position(('axes',0.2))
plt.show()
print( )

python代码中常出现的一些其他语法现象

以下这部分内容是我在暑期实习时见到及用到的一些内容,在工作中还是比较常用的,但是学校里貌似不会涉及。感兴趣的同学可以仅作了解即可,与机器学习的主题并无太大关系。

此外,由于python是解释性语言,其运行效率本身相比于C++等编译性语言就有明显的劣势,故在编程时就要尤其考虑到运行时间,即python被解释器解释后具体是如何在操作系统层面实现的,如append其实是新开辟了一块地址空间,将原先的内容拷贝过来再扩充最后一项,开辟新空间的过程耗时相对其他操作来说就是极长的。当然,这些内容是熟能生巧的,初学阶段并不需要以其为目标。

类型提示(type hints)

Python 3.5 引入了类型提示(type hints)的概念,可以为函数的参数和返回值添加类型信息,帮助编辑器提供更好的自动补全和类型检查。

类型提示可以帮助开发者更好地理解函数的用途,并在开发过程中避免一些常见的错误。

import polars as pl
from datetime import datetime
def get_contract_positions_by_day(
    product_name: str, exchange_code: str, trading_day: datetime
) -> pl.DataFrame:

以上在代码中只管展示参数类型提示和返回值类型提示,但不会影响函数的执行。这是一个用于获取每一天期货合约的函数,类型提示展示了三个参数的类型:产品名是str、股票代码是str、交易日期是datetime,也展示了这个函数返回的是一个dataframe。

“@” -- 语法糖

装饰器就是参数为函数的函数,装饰器通常用于横切关注点(cross-cutting concerns),例如日志记录、权限检查、缓存等。

def my_decorator(func):
    def wrapper():
        print("Something is happening before the function is called.")
        func()
        print("Something is happening after the function is called.")
    return wrapper

@my_decorator
def say_whee():
    print("Whee!")

say_whee()
# Output: Something is happening before the function is called. Whee! Something is happening after the function is called.

在以上代码中,

  • my_decorator 是一个装饰器函数,它接受一个函数作为参数,并返回一个新的函数。
  • say_whee 是一个函数,它被 my_decorator 装饰,因此它实际上是一个装饰过的函数。
  • 调用 say_whee() 时,实际上调用的是 wrapper 函数,它在调用 say_whee() 之前和之后打印一些日志。
  • @my_decorator 语法糖等价于 **say_whee = my_decorator(say_whee)**,即将 say_whee 赋值给 my_decorator 的返回值。
def wrapper(*args, **kwargs):
    print(f"Calling function {func.__name__} with arguments {args} and {kwargs}")

装饰器的wrapper函数接受任意数量的位置参数和关键字参数 参数解释:

  • *args:用于接收任意数量的位置参数,并将它们打包成一个元组。
  • **kwargs:用于接收任意数量的关键字参数,并将它们打包成一个字典。

lambda函数

lambda 是 Python 中的匿名函数,它是一种快速定义简单函数的方式。以下是其基本语法:

lambda 参数: 表达式

对比普通函数:

# 普通函数定义
def add(x, y):
    return x + y

# lambda 函数定义
add_lambda = lambda x, y: x + y

# 使用对比
print(add(3, 5))        # 输出: 8
print(add_lambda(3, 5)) # 输出: 8
import pandas as pd

df = pd.DataFrame({'成绩': [85, 92, 58, 88, 45]})

# 使用 lambda 判断是否及格
df['是否及格'] = df['成绩'].apply(lambda x: '是' if x >= 60 else '否')
print(df)

students = [
    {'name': '张三', 'score': 85},
    {'name': '李四', 'score': 92},
    {'name': '王五', 'score': 78}
]

# 按成绩排序
sorted_students = sorted(students, key=lambda x: x['score'], reverse=True)
print("按成绩降序:", sorted_students)

# 按姓名排序
sorted_by_name = sorted(students, key=lambda x: x['name'])
print("按姓名排序:", sorted_by_name)

“Polars” -- python库

polars是一个用于处理数据的快速、多线程 DataFrame 库。与 Pandas 相比,Polars 在性能和内存效率方面有显著优势,特别是在处理大数据集时。Polars 采用 Rust 语言编写,并在 Python 中提供了一个友好的 API,因此它可以利用 Rust 的性能优势,同时保持 Python 的易用性。

特点:

  • 高性能:利用多线程和高效的内存管理,实现了比 Pandas 更快的数据处理速度。
  • 低内存占用:优化的内存使用,使其在处理大数据集时更具优势。
  • API 易用性:提供类似于 Pandas 的 API,使得从 Pandas 迁移到 Polars 相对容易。
  • 安全性:由于使用了 Rust 语言,Polars 继承了 Rust 的内存安全特性,减少了内存泄漏和数据竞争的风险。

多进程

多进程的基本原理是在操作系统中创建多个独立的进程,每个进程拥有自己的内存空间和资源,从而能够在多个处理器核心上并行执行代码。与多线程不同,多进程不受全局解释器锁(GIL)的限制,因此在 CPU 密集型任务中能显著提高性能。

multiprocessing 模块

multiprocessing 模块提供了一种简单的方法来创建新的进程,并在这些进程之间共享数据。

import multiprocessing

def worker(num):
    print(f"Worker {num} started")
    return

if __name__ == '__main__':
    jobs = []
    for i in range(5):
        p = multiprocessing.Process(target=worker, args=(i,))
        jobs.append(p)
        p.start()
    for j in jobs:
        j.join()
  • target: 这是一个可调用对象(通常是一个函数),表示进程启动后要执行的代码。如果不指定 target,进程将不执行任何代码。
  • args: 这是一个元组,包含传递给 target 函数的位置参数。
  • target=worker 表示每个进程启动后将执行 worker 函数。
  • args=(i,) 表示传递给 worker 函数的参数是 i。
  • p.start():启动进程,执行 worker 函数。
  • p.join():阻塞主进程,直到 p 进程完成执行。
  • Output: Worker 0 started Worker 1 started Worker 2 started Worker 3 started Worker 4 started 但不一定是按顺序的,因为多进程是并行的。

ThreadPoolExecutor

  • 适用场景:I/O 密集型任务,例如网络请求、文件读写。
  • 原理:在内部创建一个线程池,任务通过线程并行执行。
  • map 方法用于将一个可调用对象和一个可迭代对象中的每个元素作为参数,并行执行该可调用对象。

类似于内置的 map 函数,但它并行地在多个线程中执行可调用对象。

from concurrent.futures import ThreadPoolExecutor

with ThreadPoolExecutor() as executor:
    results = list(
        executor.map(
            lambda x: process_symbol_id(x, all_timestamps),
            grouped_rows.values()
        )
    )

在以上代码中,会将不同的process_symbol_id任务分配到多个线程并行运行,从而提高效率。

  • 将 grouped_rows.values() 中的每个元素作为参数传递给 process_symbol_id 函数,并行执行该函数。
  • map 返回一个迭代器,生成每个函数调用的结果。
  • 匿名函数(lambda函数),可以理解为一个简洁的函数定义。它接收一个参数 x,并调用 process_symbol_id 函数,将 x 和 all_timestamps 作为参数传递。

ProcessPoolExecutor

  • 适用场景:CPU 密集型任务,例如图像处理、机器学习。
  • 原理:在内部创建一个进程池,任务通过进程并行执行。
  • submit: 这个方法将一个可调用对象及其参数提交给进程池执行。它立即返回一个 Future 对象,表示待执行或正在执行的操作。
  • fn: 可调用对象,即要执行的函数。
  • *args: 传递给 fn 的位置参数。
  • **kwargs: 传递给 fn 的关键字参数。
from concurrent.futures import ProcessPoolExecutor

with ProcessPoolExecutor() as executor:
    futures = [
        executor.submit(process_trading_day, product_name, exchange_code, day)
        for day in trading_days
    ]

在以上代码中,会将同一个id的内容按不同日期并发执行。 

  • executor.submit(process_trading_day, product_name, exchange_code, day) 将 process_trading_day 函数及其参数提交给进程池执行。
  • submit 返回一个 Future 对象,用于获取任务的结果。

异步处理

Python 提供了 asyncio 库,支持异步编程,使用 async 定义异步函数,使用 await 等待异步操作的完成。

  • 同步:任务按顺序执行,一个任务完成后才会执行下一个任务。如果某个任务需要较长时间(如网络请求),整个程序会在这段时间内被阻塞,直到任务完成。
  • 异步:任务可以被分解并并行处理,允许程序在等待某些任务完成的同时继续执行其他任务,从而避免了阻塞。

异步应用的特点:

  1. 非阻塞:当程序执行一个需要等待的操作(如网络请求)时,程序不会停止等待结果,而是继续执行其他任务。操作完成后,程序会收到通知并处理结果。
  2. 并发执行:在异步模型中,多个任务可以在同一时间段内执行(不是同时启动,而是重叠执行),提高了资源利用率和程序效率。
  3. 回调机制:异步操作通常通过回调函数或事件循环的方式来处理完成后的任务。Python 中使用 async 和 await 关键字来简化异步代码的编写。回调机制在处理与应用程序、用户、或随着时间变动的内容交互时还是挺常见的。
import asyncio
async def fetch_data():
    print("Fetching data...")
    await asyncio.sleep(2)  # 模拟耗时操作
    print("Data fetched")
async def main():
    await fetch_data()
    print("Main function continues...")
# 运行异步应用
asyncio.run(main())

gc.collect() -- 释放内存的函数

Python 的垃圾回收器会自动回收不再使用的内存,但有时需要手动调用 gc.collect() 来释放内存。当内存占用过高时,可以调用 gc.collect() 来释放内存,尤其是对于中间变量,但这可能会导致程序变慢。

建议:在程序中尽量避免频繁调用 gc.collect(),因为它会影响程序的运行速度。

import gc
obj = a+b
c = obj+b
del obj  # 手动删除对象
gc.collect()  # 手动触发垃圾回收 (尤其是循环引用的变量)

总结

本篇文章总结了Python在机器学习中的基础知识,介绍了Python的灵活性和广泛应用,强调了列表、元组、字典、集合、字符串等常用数据结构的特点和用法,并简要提到了开发环境的选择与安装库的建议。总体上为初学者提供了一个简洁明了的Python入门指南,为后续学习机器学习打下基础。

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐