日常 Python 操作经常围绕文本、索引和循环展开。先确认数据类型,再区分“生成一个值”“修改对象”和“重复执行”,可以避开许多看似相近的写法。本文以 Python 3.11 为基线;最后的图像裁剪例子额外使用 NumPy。
数字文本、补零、清理与字典更新
跳转到“数字文本、补零、清理与字典更新”int("000020") 解析十进制整数,str(20).zfill(6) 生成带前导零的文本。两者转换方向相反。strip("0") 则从两端移除字符 0,既不是数字转换,也不是只删除前导零:对原例 "000020",结果为 "2"。
import math
num_str = "000020"num = int(num_str)assert num == 20assert str(num).zfill(6) == "000020"assert num_str.strip("0") == "2"assert num_str.lstrip("0") == "20"assert num_str.rstrip("0") == "00002"assert "000".lstrip("0") == "" and int("000") == 0assert " hello world! \n".strip() == "hello world!"assert "a-b-a".replace("a", "x") == "x-b-x"
# 原记录中的两项数值没有注明单位,这里只演示十进制文本转换。raw_values = ["113.8029", "116.3988"]values = [float(value) for value in raw_values]assert math.isclose(values[1] - values[0], 2.5959)
dict1 = {"a": 1, "b": 2}dict2 = {"c": 3, "d": 4}result = dict1.update(dict2)assert result is Noneassert dict1 == {"a": 1, "b": 2, "c": 3, "d": 4}dict1.update({"a": 9})assert dict1["a"] == 9print(num, str(num).zfill(6), dict1)strip() 默认处理空格、制表符、换行等空白字符,不是只删除 ASCII 空格。lstrip()、rstrip() 分别处理左、右两端;replace() 替换匹配子串。处理数字时优先按数字格式解析,不要先随意删字符掩盖错误。方法边界见字符串,字典更新与复制见字典。
输出在一行、同时取得索引和值
跳转到“输出在一行、同时取得索引和值”print 默认以换行结束,end=" " 改为一个空格;sep 则控制同一次调用中多个参数之间的分隔。下面把输出暂存到 StringIO,使空格和换行也能被准确检查。
from io import StringIO
output = StringIO()for i in range(10): print(i, end=" ", file=output)assert output.getvalue() == "0 1 2 3 4 5 6 7 8 9 "print(output.getvalue().rstrip())
colors = ["red", "green", "blue"]pairs = list(enumerate(colors))assert pairs == [(0, "red"), (1, "green"), (2, "blue")]for index, color in pairs: print(index, color)assert list(enumerate(colors, start=1))[0] == (1, "red")
linemod_cls_names = ["ape", "benchvise", "camera"]class_names = {cls_id: cls_name for cls_id, cls_name in enumerate(linemod_cls_names)}assert class_names == {0: "ape", 1: "benchvise", 2: "camera"}enumerate 产生计数与元素;名称写成 cls_id、cls_name 不会自动决定真实数据集的类别编号。这里仅展示原笔记的变量用途,示例列表是本程序定义的三项列表。若正式数据另有编号规则,应使用该规则,而不是直接假定从零连续编号。
切片:终点不包含,方向由步长决定
跳转到“切片:终点不包含,方向由步长决定”通用写法为 sequence[start:stop:step]。第二个冒号用于引入步长;不是只有写成相邻的 :: 才能指定步长,a[1:8:2] 同样合法。步长不能为零。
正步长时省略的边界覆盖从左到右的范围,负步长时省略的边界适配从右到左。负索引是相对结尾的位置;负步长才决定反向遍历。不要把二者混为一谈,也不要把负步长的终点简单读成“到 stop-1”。
a = list(range(12))assert a[1:5] == [1, 2, 3, 4]assert a[:5] == [0, 1, 2, 3, 4]assert a[3:] == list(range(3, 12))assert a[1:5:2] == [1, 3]assert a[1:8:2] == [1, 3, 5, 7]assert a[::2] == [0, 2, 4, 6, 8, 10]assert a[::3] == [0, 3, 6, 9]assert a[::-1] == list(reversed(a))assert a[-5:-2] == [7, 8, 9]assert a[-6:10:2] == [6, 8]tmp = [1, 2, 3, 4, 5, 6]assert tmp[5::-2] == [6, 4, 2]assert tmp[5:-1:-2] == []assert a[100:200] == []try: a[::0]except ValueError: passelse: raise AssertionError("slice step cannot be zero")
original = [[1], [2]]shallow = original[:]assert shallow == original and shallow is not originalassert shallow[0] is original[0]assert "abcdef"[::2] == "ace"assert (1, 2, 3)[::-1] == (3, 2, 1)print("slice checks passed")tmp[5::-2] 省略了终点;tmp[5:-1:-2] 则明确把终点设为最后一个元素的位置,因此该例为空。切片越界通常被裁剪,而单个越界索引会抛出 IndexError。
列表的完整切片生成浅复制;嵌套成员仍可共享。不能从 sequence[:] 推出所有类型都会创建一个独立对象:不可变序列可能复用对象,NumPy 数组的基础切片通常是视图。详见列表复制关系和 Python 序列切片。
for、while 与 range 的实际执行次数
跳转到“for、while 与 range 的实际执行次数”for 从可迭代对象逐项获取值,不是统一实现成“先赋值再比较上界”。可迭代对象可能没有 len(),也可能无限产生值;break、异常或迭代期间的数据变化都会影响循环次数。while 则在每轮开始前判断条件,通常需要明确地更新使条件最终为假的状态。
下面同时保留 range(1, 4)、步长为一的 range(2, 129) 和步长为二的 range(2, 129, 2):
seen = []for item in [1, 2, 3, 4, 5]: seen.append(item)assert seen == [1, 2, 3, 4, 5]
count = 0while count < 5: count += 1assert count == 5
for args, expected_count, expected_last in [ ((1, 4), 3, 3), ((2, 129), 127, 128), ((2, 129, 2), 64, 128), ((5, -2, -2), 4, -1),]: sequence = range(*args) visits = 0 for i in sequence: visits += 1 assert visits == len(sequence) == expected_count assert i == sequence[-1] == expected_last
visits = 0for item in (n * n for n in range(100)): visits += 1 if visits == 3: breakassert visits == 3
i = "unchanged"for i in range(3, 3): raise AssertionError("empty range cannot enter loop")assert i == "unchanged"print("loop checks passed")空循环不会为循环变量产生一个“最后的值”:若变量此前存在,它保持原值;若此前未绑定,之后直接读取会报错。非空、完整执行的 range(start, stop, step) 才可以按以下计数和最后一项公式分析。
用整数公式计算范围与数列
跳转到“用整数公式计算范围与数列”令起点为 、终点为 、步长为 。正步长要求每项小于终点,负步长要求每项大于终点;方向不匹配时范围为空。若范围非空,项数和最后一项为:
这个公式的非空前提不能省略。实际程序可直接使用 len(range(...));对可能超过平台长度上限的大范围,可用纯整数计算项数,避免先做浮点除法再取整而丢失精度。
等差数列的第 项及前 项和为:
等比数列使用公比 :
这些项数公式针对正整数 。两类数列都是离散序列;不能把“等差数列的连续项”解释成连续变量。循环是执行机制,可以生成数列,也可以执行与数列无关的操作。
def range_count(start, stop, step=1): if step == 0: raise ValueError("step must not be zero") if step > 0: return max(0, (stop - start + step - 1) // step) size = -step return max(0, (start - stop + size - 1) // size)
for start in range(-4, 5): for stop in range(-4, 5): for step in [-3, -2, -1, 1, 2, 3]: actual = range(start, stop, step) n = range_count(start, stop, step) assert n == len(actual) if n: assert start + (n - 1) * step == actual[-1]assert range_count(2, 129, 2) == 64assert range_count(0, 10**40, 3) == (10**40 + 2) // 3
arithmetic = list(range(2, 129, 2))n, first, last = len(arithmetic), arithmetic[0], arithmetic[-1]assert n == 64 and last == first + (n - 1) * 2assert sum(arithmetic) == n * (first + last) // 2 == 4160
first, ratio, n = 3, 2, 4geometric = [first * ratio**k for k in range(n)]assert geometric == [3, 6, 12, 24]assert geometric[-1] == first * ratio ** (n - 1)assert sum(geometric) == first * (ratio**n - 1) // (ratio - 1) == 45assert sum([3] * 4) == 4 * 3print("range and sequence formula checks passed")range_count 示例约定传入整数;这里的整数除法适用于示例中的整数组合,不能把等比求和公式里的除号在任意小数公比下都换成 //。len(range(...)) 也可能因长度超出平台可表示范围而报 OverflowError,这与范围本身无法表示是两件事。
向上、向下、截断与就近取整
跳转到“向上、向下、截断与就近取整”math.ceil 向正无穷方向、math.floor 向负无穷方向,math.trunc 和有限浮点数的 int 转换向零截断。Python 内置 round 在精确中点采用“取偶数”,不是一律把 .5 远离零。
import mathfrom decimal import Decimal, ROUND_HALF_UP
assert math.ceil(3.1) == 4 and math.ceil(-3.7) == -3assert math.floor(3.7) == 3 and math.floor(-3.1) == -4assert math.trunc(3.9) == 3 and math.trunc(-3.9) == -3assert int(3.9) == 3 and int(-3.9) == -3assert round(3.5) == 4 and round(2.5) == 2assert round(-3.5) == -4 and round(-2.5) == -2assert round(2.4) == 2assert round(2.675, 2) == 2.67assert Decimal("2.5").quantize(Decimal("1"), rounding=ROUND_HALF_UP) == Decimal("3")assert Decimal("-2.5").quantize(Decimal("1"), rounding=ROUND_HALF_UP) == Decimal("-3")print("rounding checks passed")2.675 的例子还涉及二进制浮点近似,并非只看十进制书写形式就能决定结果。若需求明确规定十进制舍入规则,用字符串构造 Decimal 并选择相应模式。依据:round、math 的取整函数。
类型注解与一个完整的裁剪函数
跳转到“类型注解与一个完整的裁剪函数”函数注解语法和类型提示的发展有不同时间节点:函数注解在 Python 3.0 已有,标准类型提示方案见 Python 3.5 的 PEP 484,value: int = 5 这种变量注解语法由 Python 3.6 引入。注解本身不会强制运行时检查,原稿只写 pass 的 image_crop 也不会返回裁剪图像。
下面将边界明确为 (x0, y0, x1, y1),右端和下端不包含;要求非空且完全位于图像内,返回独立副本。代码使用 NumPy,可处理二维灰度图或带额外通道维度的数组。
from collections.abc import Sequenceimport operatorimport numpy as np
def image_crop(img: np.ndarray, bbox: Sequence[int]) -> np.ndarray: if not isinstance(img, np.ndarray) or img.ndim < 2: raise ValueError("img must be an array with at least two dimensions") if len(bbox) != 4: raise ValueError("bbox must contain x0, y0, x1, y1") if any(isinstance(value, (bool, np.bool_)) for value in bbox): raise TypeError("boolean coordinates are not accepted") x0, y0, x1, y1 = (operator.index(value) for value in bbox) height, width = img.shape[:2] if not (0 <= x0 < x1 <= width and 0 <= y0 < y1 <= height): raise ValueError("bbox must be nonempty and inside the image") return img[y0:y1, x0:x1].copy()
value_name: int = 5assert value_name == 5image = np.arange(4 * 5 * 3).reshape(4, 5, 3)cropped = image_crop(image, [1, 1, 4, 3])assert cropped.shape == (2, 3, 3)assert np.array_equal(cropped, image[1:3, 1:4])assert not np.shares_memory(cropped, image)cropped[0, 0, 0] = -1assert image[1, 1, 0] != -1try: image_crop(image, [0, 0, 6, 3])except ValueError: passelse: raise AssertionError("out-of-bounds crop must be rejected")print("annotated image crop checks passed")这是为原接口补充的一种明确实现;需要裁剪越界框、归一化坐标或保留共享视图的程序应另定规则。原稿导入的 Union 没有在签名中使用,因而此例不导入它。注解详见类型注解,数组视图与复制见 NumPy 基础索引。
读取文本:保留哪部分空白要先决定
跳转到“读取文本:保留哪部分空白要先决定”原笔记的“txt 读取”只有空代码块。下面补充一个完整示例:在临时目录创建 UTF-8 文件,然后逐行读取;离开临时目录上下文后清理这个演示文件。strip() 会删除两端所有空白,若缩进或尾部空格有意义,就只去掉行结束符。
from pathlib import Pathfrom tempfile import TemporaryDirectory
with TemporaryDirectory() as directory: filename = Path(directory) / "sample.txt" filename.write_text(" alpha \nbeta\n\n", encoding="utf-8") with filename.open("r", encoding="utf-8") as stream: rows = [line.rstrip("\r\n") for line in stream] assert rows == [" alpha ", "beta", ""] assert [row.strip() for row in rows] == ["alpha", "beta", ""] assert " a b ".split(" ") == ["", "a", "", "b", ""] assert " a b\t".split() == ["a", "b"]print("text file checks passed")split(" ") 只把单个空格当分隔符并保留相邻分隔造成的空字段;不传参数的 split() 才按连续空白分组。更多输入输出参数见 print、enumerate、range与 open。